Data processing method and apparatus, electronic device, and medium

By performing similarity analysis on the speech data, qualified speech data is selected, which solves the problem of unqualified speech data affecting model training and improves the efficiency and accuracy of model training.

CN114187924BActive Publication Date: 2025-10-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111490985.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2025-10-28
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

In existing technologies, there are unqualified data in speech data, which leads to slower model training convergence speed and reduced accuracy, and there is a lack of effective screening methods.

Method used

By performing similarity analysis on multiple voice data, the similarity value between the voice data and other voice data is used to determine whether the voice data is qualified data.

Benefits of technology

This allows for the rapid and effective selection of qualified speech data that reflects the characteristics of a subject's speech, thereby improving the efficiency and accuracy of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187924B_ABST
    Figure CN114187924B_ABST
Patent Text Reader

Abstract

This disclosure provides a data processing method, apparatus, electronic device, and medium, relating to the field of artificial intelligence, and particularly to the field of speech technology. The implementation involves: determining multiple speech data associated with a first object, wherein each of the multiple speech data has a tag for identifying the first object; and for any one of the multiple speech data, determining whether the speech data is qualified speech data for the first object based on the similarity value between the speech data and the other speech data in the multiple speech data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the field of voice technology, specifically to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for data processing.

[0005] According to one aspect of this disclosure, a data processing method is provided, comprising: determining a plurality of voice data associated with a first object, wherein each of the plurality of voice data has a label for identifying the first object; and determining, for any one of the plurality of voice data, whether the voice data is qualified voice data of the first object based on a similarity value between the voice data and other voice data in the plurality of voice data.

[0006] According to another aspect of this disclosure, a data processing apparatus is provided, comprising: a first determining unit configured to determine a plurality of voice data associated with a first object, wherein each of the plurality of voice data has a tag for identifying the first object; and a second determining unit configured to, for any one of the plurality of voice data, determine whether the voice data is qualified voice data of the first object based on a similarity value between the voice data and other voice data in the plurality of voice data.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described above.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described method when executed by a processor.

[0010] According to one or more embodiments of this disclosure, qualified voice data that can reflect the voice characteristics of an object can be quickly and effectively filtered out.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0013] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;

[0014] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown;

[0015] Figure 3 A flowchart of another data processing method according to an embodiment of the present disclosure is shown;

[0016] Figure 4 A structural block diagram of a data processing apparatus according to an embodiment of the present disclosure is shown;

[0017] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0018] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0019] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0020] The terminology used in the description of the various examples in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0021] To support the needs of applications such as model training, a large amount of high-quality data is required. For example, in the field of speech synthesis technology, in order to train a speech synthesis model for a specific object so that the trained model can automatically synthesize synthesized speech that conforms to the vocal characteristics of that specific object, a large amount of speech data that accurately reflects the vocal characteristics of that specific object is needed. However, in real-world scenarios, due to the influence of environmental factors, equipment, or operational errors, the obtained speech data often contains a small number of unqualified speech data, such as speech data with incorrect labels or speech data with significant environmental noise. Applying such unqualified speech data to model training will slow down the convergence speed of model training and reduce model accuracy.

[0022] In related technologies, the quality of voice data is determined by comparing it with known standard data. However, there is no effective voice screening method when standard data is unavailable.

[0023] Based on this, this disclosure proposes a data processing method that, for each voice data in a plurality of voice data sets, determines whether the voice data is qualified based on its similarity value with each of the other voice data in the plurality of voice data sets. Thus, it is possible to eliminate a small number of unqualified voice data sets from a plurality of voice data sets that are affected by environmental factors, equipment, etc., and quickly and effectively filter out qualified voice data that can reflect the voice characteristics of the first object.

[0024] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0025] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0026] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of data processing methods.

[0027] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0028] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0029] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to acquire multiple voice data points. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0030] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0031] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0032] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0033] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0034] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.

[0035] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0036] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0037] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0038] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0039] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0040] Figure 2 An exemplary embodiment of the present disclosure illustrates a data processing method, comprising: step S201, determining a plurality of voice data associated with a first object, wherein each of the plurality of voice data has a tag for identifying the first object; and step S202, for any one of the plurality of voice data, determining whether the voice data is qualified voice data of the first object based on a similarity value between the voice data and other voice data in the plurality of voice data.

[0041] Therefore, based on the similarity value between each voice data and each of the other voice data in the multiple voice data, it is possible to identify a small number of unqualified voice data that are affected by the environment, equipment, etc., and then quickly and effectively filter out qualified voice data that can reflect the voice characteristics of the first object.

[0042] Regarding step S201, according to some embodiments, the first object can be a character in the audio text. The audio text includes corresponding text data and audio data; each sentence in the audio text is accompanied by corresponding voice data. For the dialogue of different characters in the audio text, different objects will perform expressive and engaging acts, incorporating an understanding of the character's personality, age, and other characteristics. This is reflected in the auditory experience by adjusting different voices for different characters.

[0043] According to some embodiments, multiple voice data associated with the first object may be obtained by extracting from the entire audio data of the spoken text.

[0044] According to some embodiments, the spoken text includes at least multiple audio dialogue segments. The method may further include: determining a role label for each audio dialogue segment, and wherein determining multiple voice data associated with a first object may include: for each audio dialogue segment, in response to the role label of that audio dialogue segment being a first object, determining that audio dialogue segment as voice data associated with the first object. Thus, based on the role label of each audio dialogue segment, each audio dialogue segment can be easily associated with a corresponding role, facilitating the aggregation, organization, and correction of the audio dialogue for each role.

[0045] The multiple audio dialogue segments can be all the audio dialogue in the audio text, or only part of the audio dialogue in the audio text; this disclosure does not limit this.

[0046] According to some embodiments, the audio text may further include multiple dialogue texts corresponding to multiple audio segments, and as follows: Figure 3 As shown, determining the role label of each dialogue audio segment in multiple dialogue audio segments can include: step S301, performing character recognition on the text segment containing each dialogue text in the audio text to obtain the recognition result for that dialogue text; and step S302, determining the role label of the corresponding dialogue audio segment based on the recognition result of each dialogue text in the multiple dialogue text segments. Therefore, based on the correspondence between text and audio in the audio text, the role label of the corresponding dialogue audio segment can be determined through character recognition of the dialogue text.

[0047] For example, the audio text data includes the following paragraph: "Zhang San said unhappily, 'Leave your money before you leave.'" For the dialogue text "Leave your money before you leave," the entire text paragraph containing this dialogue, namely "Zhang San said unhappily, 'Leave your money before you leave,'" is subjected to character recognition to obtain the recognition result for this dialogue text, namely, the character name "Zhang San" is identified. Based on the identified character name "Zhang San," the character tag corresponding to the audio dialogue text can be determined to be "Zhang San."

[0048] According to some embodiments, determining the role label of each dialogue audio segment in a plurality of dialogue audio segments may include: determining the role label of each dialogue audio segment in a plurality of dialogue audio segments using a trained speech recognition model. This allows for convenient determination of the role labels of dialogue audio segments using a trained model.

[0049] Regarding step S202, according to some embodiments, for each of the multiple voice data, it can be determined whether the voice data is a qualified voice data of the first object based on the similarity value between the voice data and each of the other voice data in the multiple voice data.

[0050] According to some embodiments, determining whether a voice data is qualified voice data of a first object based on the similarity value between the voice data and other voice data in a plurality of voice data may include: determining the number of voice data in other voice data whose similarity value with the voice data is higher than a preset threshold; and determining whether the voice data is qualified voice data of the first object based on the number.

[0051] The more similarity data a given voice data point has to a preset threshold, the higher its matching degree with the first object, meaning the more representative the voice data is of the first object's voice characteristics. Conversely, if the similarity value is lower, the voice data fails to reflect the first object's voice characteristics. Therefore, by determining the number of voice data points with similarity values ​​exceeding the preset threshold, the quality of the voice data can be quantified, thus facilitating the determination of whether the voice data is qualified for the first object.

[0052] Among them, the other voice data in the multiple voice data can be all voice data other than the voice data in the multiple voice data, or it can be a part of the voice data other than the voice data in the multiple voice data.

[0053] According to some embodiments, the method further includes: selecting a portion of voice data from all voice data other than the voice data in a plurality of voice data before determining whether the voice data is qualified voice data for the first object; and wherein determining whether the voice data is qualified voice data for the first object based on the similarity value between the voice data and each of the other voice data in the plurality of voice data may include: determining whether the voice data is qualified voice data for the first object based on the similarity value between the voice data and each of the selected portion of voice data.

[0054] According to some embodiments, a portion of the voice data selected from multiple voice data can be randomly selected from multiple voice data based on a preset number.

[0055] According to some embodiments, the similarity value includes a timbre similarity value.

[0056] Figure 4 A data processing apparatus 400 according to an exemplary embodiment of the present disclosure is shown. The apparatus 400 includes: a first determining unit 401 configured to determine a plurality of voice data associated with a first object, wherein each of the plurality of voice data has a tag for identifying the first object; and a second determining unit 402 configured to determine, for any one of the plurality of voice data, whether the voice data is qualified voice data of the first object based on a similarity value between the voice data and other voice data in the plurality of voice data.

[0057] According to some embodiments, the second determining unit includes: a subunit for determining the number of voice data in other voice data whose similarity value to the voice data is higher than a preset threshold; and a subunit for determining, based on the number, whether the voice data is qualified voice data of the first object.

[0058] According to some embodiments, the similarity value includes a timbre similarity value.

[0059] According to some embodiments, the first object is a character in an audio text.

[0060] According to some embodiments, the spoken text includes at least multiple dialogue audio segments, and the apparatus further includes: a third determining unit configured to determine a role label for each of the multiple dialogue audio segments, and wherein the first determining unit further includes: for each of the multiple dialogue audio segments, in response to the role label of the dialogue audio segment being a first object, determining the dialogue audio segment as a subunit of voice data associated with the first object.

[0061] According to some embodiments, the audio text also includes multiple dialogue texts corresponding to multiple dialogue audio segments, and wherein the third determining unit includes: a subunit for performing character recognition on the text segment in the audio text containing each dialogue text in the multiple dialogue texts to obtain a recognition result for the dialogue text; and a subunit for determining the role tag of the dialogue audio corresponding to the dialogue text based on the recognition result of each dialogue text in the multiple dialogue texts.

[0062] According to some embodiments, the third determining unit includes a subunit for determining the role label of each dialogue audio segment in a plurality of dialogue audio segments by means of a trained speech recognition model.

[0063] According to embodiments of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described above.

[0064] According to embodiments of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause a computer to perform any of the methods described above.

[0065] According to embodiments of this disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements any of the methods described above when executed by a processor.

[0066] refer to Figure 5 The present invention describes a structural block diagram of an electronic device 500 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0067] like Figure 5As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0068] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0069] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform data processing methods by any other suitable means (e.g., by means of firmware).

[0070] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0071] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0072] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0073] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0074] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0075] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0076] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0077] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A data processing method, comprising: A plurality of voice data associated with a first object are determined, wherein each of the plurality of voice data is dialogue audio with a tag for identifying the first object; For each of the plurality of voice data: Based on the similarity value between the voice data and other voice data in the plurality of voice data, it is determined whether the voice data is qualified voice data for the first object. The determination of whether the voice data is qualified voice data for the first object based on the similarity value between the voice data and other voice data in the plurality of voice data includes: Determine the number of voice data points in the other voice data whose similarity value to the given voice data is higher than a preset threshold; and Based on the quantity, determine whether the voice data is qualified voice data for the first object; and In response to determining, based on the quantity, that the voice data is invalid due to environmental noise interference, and delete the voice data from the plurality of voice data; and Multiple voice data points are obtained after deleting unqualified voice data, and used as training samples. The training samples are used to train a voice synthesis model for automatically synthesizing synthesized voice that conforms to the vocal characteristics of the first object.

2. The method according to claim 1, wherein, The similarity value includes timbre similarity value.

3. The method according to claim 1 or 2, wherein, The first object is a character in an audio text.

4. The method according to claim 3, wherein, The spoken text includes at least multiple audio dialogue segments, and the method further includes: Determine the role label for each of the multiple audio dialogue segments. Furthermore, the determination of the plurality of voice data associated with the first object includes: For each of the multiple dialogue audio segments, in response to the role label of the dialogue audio segment being a first object, the dialogue audio segment is determined as voice data associated with the first object.

5. The method according to claim 4, wherein, The audio text also includes multiple dialogue texts corresponding to the multiple audio dialogue segments, and wherein determining the role label for each of the multiple audio dialogue segments includes: For each of the plurality of dialogue texts, character recognition is performed on the text segment containing that dialogue text in the audio text to obtain the recognition result for that dialogue text; and Based on the recognition result of each of the plurality of dialogue texts, the role label of the dialogue audio corresponding to that dialogue text is determined.

6. The method according to claim 4, wherein, Determining the role label for each segment of dialogue audio in the multiple audio segments includes: The role label for each segment of dialogue audio is determined using a trained speech recognition model.

7. A data processing apparatus, comprising: A first determining unit is configured to determine a plurality of voice data associated with a first object, wherein each of the plurality of voice data is dialogue audio with a tag for identifying the first object; and The second determining unit is configured to determine each of the plurality of voice data: Based on the similarity value between the voice data and other voice data in the plurality of voice data, it is determined whether the voice data is qualified voice data for the first object. The determination of whether the voice data is qualified voice data for the first object based on the similarity value between the voice data and other voice data in the plurality of voice data includes: Determine the number of voice data points in the other voice data whose similarity value to the given voice data is higher than a preset threshold; and Based on the quantity, determine whether the voice data is qualified voice data for the first object; and In response to determining, based on the quantity, that the voice data is invalid due to environmental noise interference, and delete the voice data from the plurality of voice data; and Units for acquiring multiple voice data after deleting unqualified voice data, which are used as training samples, are used to train a voice synthesis model for automatically synthesizing synthesized voice that conforms to the vocal characteristics of the first object.

8. The apparatus according to claim 7, wherein, The similarity value includes timbre similarity value.

9. The apparatus according to claim 7 or 8, wherein, The first object is a character in an audio text.

10. The apparatus according to claim 9, wherein, The spoken text includes at least multiple audio dialogue segments, and the device further includes: The third determining unit is configured to determine the role label for each of the multiple dialogue audio segments. Furthermore, the first determining unit further includes: For each of the multiple dialogue audio segments, in response to the role label of the dialogue audio segment being a first object, the dialogue audio segment is identified as a subunit of voice data associated with the first object.

11. The apparatus according to claim 10, wherein, The audio text also includes multiple dialogue texts corresponding to the multiple audio segments, and wherein the third determining unit includes: A subunit for performing character recognition on the text segment containing each of the plurality of dialogue texts in the audio text, to obtain the recognition result for that dialogue text; and A subunit used to determine the role label of the dialogue audio corresponding to each of the plurality of dialogue texts based on the recognition result of each dialogue text.

12. The apparatus according to claim 10, wherein, The third determining unit includes: A subunit used to determine the role label of each segment of dialogue audio in the multi-segment dialogue audio using a trained speech recognition model.

13. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Speech data processing method, device and equipment, and storage medium

    CN109616097A

  • Article voice playing method, apparatus and device, and computer readable storage medium

    CN113010138A