Speech separation method
By employing a gating processing method based on local and global attention mechanisms, the problem of poor speech separation in multi-person communication is solved, achieving effective speech separation and recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies cannot effectively separate speech during multi-person communication, affecting speech recognition systems and auditory experience, resulting in poor speech separation performance.
A gating processing method employing local and global attention mechanisms is used to extract speech features from speech information sequences and separate speech information from different pronunciation objects through speech mask information.
It achieves effective separation of speech, reduces the computational requirements of the attention mechanism, can process global and local speech information, and improves the accuracy and auditory experience of the speech recognition system.
Smart Images

Figure CN116168717B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio processing, and in particular, to a speech separation method. BACKGROUND
[0002] At present, if speech separation is not performed when multiple people communicate at the same time, it will directly affect the speech recognition system or the hearing perception and understanding degree.
[0003] In the related art, when processing speech, only a single source speech is directly separated from a single overlapping mixed speech. Since this method is to extract mixed speech directly for speech separation, there is a technical problem of poor effect of speech separation on speech.
[0004] In view of the above problems, no effective solution has been proposed so far. SUMMARY
[0005] The embodiments of the present application provide a speech separation method to at least solve the technical problem that speech cannot be separated.
[0006] According to an aspect of an embodiment of the present application, a speech separation method is provided. The method can include: obtaining a speech information sequence, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different sound producing objects; extracting speech features of different sound producing objects from the speech information sequence to obtain a speech feature sequence; performing gate processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gate processing result, wherein the gate processing result includes local speech information and global speech information of different sound producing objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound producing objects based on the gate processing result, wherein the speech mask information is used to represent the sound producing properties of the sound producing objects; and separating the speech information output by different sound producing objects from the speech information sequence based on the speech mask information of different sound producing objects and the speech feature sequence.
[0007] According to another aspect of the embodiments of the present application, another speech separation method is provided. The method can include: obtaining a speech information sequence, wherein the speech information sequence comprises at least one speech information to be separated, and different speech information comes from different speech objects; calling a speech separation model, wherein the speech separation model is obtained by training based on a local attention mechanism and a global attention mechanism; using the speech separation model to extract speech features of different speech objects from the speech information sequence to obtain a speech feature sequence, and performing gating processing on the speech features in the speech feature sequence according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different speech objects based on the gating processing result, wherein the speech mask information is used to represent speech attributes of the speech objects; and separating speech information output by different speech objects from the speech information sequence based on the speech mask information of different speech objects and the speech feature sequence.
[0008] According to another aspect of the embodiments of the present application, another speech separation method is provided. The method can include: obtaining a speech information sequence, wherein the speech information sequence comprises at least one speech information to be separated, and different speech information comes from different speech objects; calling a speech separation model, wherein the speech separation model is obtained by training based on a local attention mechanism and a global attention mechanism; using the speech separation model to extract speech features of different speech objects from the speech information sequence to obtain a speech feature sequence, and performing gating processing on the speech features in the speech feature sequence according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different speech objects based on the gating processing result, wherein the speech mask information is used to represent speech attributes of the speech objects; and separating speech information output by different speech objects from the speech information sequence based on the speech mask information of different speech objects and the speech feature sequence.
[0009] According to another aspect of the embodiments of the present application, another speech separation method is provided. The method can include: extracting speech features of different speech objects from an acquired speech information sequence to obtain a speech feature sequence, wherein the speech information sequence comprises at least one speech information to be subjected to speech separation, and the different speech information is from different speech objects; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of the different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of the different speech objects based on the gating processing result, wherein the speech mask information is used to represent speech attributes of the speech objects; separating speech information output by the different speech objects from the speech information sequence based on the speech mask information of the different speech objects and the speech feature sequence; and inputting the speech information output by the different speech objects to a speech recognition end, wherein the speech information is used to be recognized by the speech recognition end.
[0010] According to another aspect of the embodiments of the present application, another speech separation method is provided. The method can include: extracting speech features of different speech objects from an acquired speech information sequence to obtain a speech feature sequence, wherein the speech information sequence comprises at least one speech information to be subjected to speech separation, and the different speech information is from different speech objects; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of the different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of the different speech objects based on the gating processing result, wherein the speech mask information is used to represent speech attributes of the speech objects; separating speech information output by the different speech objects from the speech information sequence based on the speech mask information of the different speech objects and the speech feature sequence; and inputting the speech information output by the different speech objects to a speech recognition end, wherein the speech information is used to be recognized by the speech recognition end.
[0011] According to another aspect of the embodiments of the present application, a speech separation device is provided. The device can include: a first obtaining unit configured to obtain a sequence of speech information, wherein the sequence of speech information comprises at least one speech information to be separated, and different speech information is from different speech objects; a first extracting unit configured to extract speech features of different speech objects from the sequence of speech information to obtain a sequence of speech features; a first processing unit configured to perform gating processing on the speech features in the sequence of speech features according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; a second obtaining unit configured to obtain speech mask information of different speech objects based on the gating processing result, wherein the speech mask information is used to represent speech properties of the speech objects; and a first separating unit configured to separate speech information output by different speech objects from the sequence of speech information based on the speech mask information of different speech objects and the sequence of speech features.
[0012] According to another aspect of the embodiments of the present application, another speech separation device is provided. The device can include: a third obtaining unit configured to obtain a sequence of speech information, wherein the sequence of speech information comprises at least one speech information to be separated, and different speech information is from different speech objects; a first calling unit configured to call a speech separation model, wherein the speech separation model is obtained by training based on a local attention mechanism and a global attention mechanism; a second extracting unit configured to extract speech features of different speech objects from the sequence of speech information using the speech separation model to obtain a sequence of speech features, and perform gating processing on the speech features in the sequence of speech features according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; a fourth obtaining unit configured to obtain speech mask information of different speech objects based on the gating processing result, wherein the speech mask information is used to represent speech properties of the speech objects; and a second separating unit configured to separate speech information output by different speech objects from the sequence of speech information based on the speech mask information of different speech objects and the sequence of speech features.
[0013] According to another aspect of the embodiments of the present application, another voice separation device is provided. The device can include: a third extraction unit configured to extract speech features of different utterance objects from a sequence of acquired speech information to obtain a sequence of speech features, wherein the sequence of speech information comprises at least one speech information to be subjected to voice separation, and the different speech information is from different utterance objects; a second processing unit configured to perform gating processing on the speech features in the sequence of speech features according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of the different utterance objects, and the information granularity of the local speech information is smaller than that of the global speech information; a fifth acquisition unit configured to acquire speech mask information of the different utterance objects based on the gating processing result, wherein the speech mask information is used to represent utterance attributes of the utterance objects; a third separation unit configured to separate speech information output by the different utterance objects from the sequence of speech information based on the speech mask information of the different utterance objects and the sequence of speech features; and a playing unit configured to play the speech information output by the different utterance objects respectively.
[0014] According to another aspect of the embodiments of the present application, another voice separation device is provided. The device can include: a fourth extraction unit configured to extract speech features of different utterance objects from a sequence of acquired speech information to obtain a sequence of speech features, wherein the sequence of speech information comprises at least one speech information to be subjected to voice separation, and the different speech information is from different utterance objects; a third processing unit configured to perform gating processing on the speech features in the sequence of speech features according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of the different utterance objects, and the information granularity of the local speech information is smaller than that of the global speech information; a sixth acquisition unit configured to acquire speech mask information of the different utterance objects based on the gating processing result, wherein the speech mask information is used to represent utterance attributes of the utterance objects; a fourth separation unit configured to separate speech information output by the different utterance objects from the sequence of speech information based on the speech mask information of the different utterance objects and the sequence of speech features; and an input unit configured to input the speech information output by the different utterance objects to a speech recognition end, wherein the speech information is used to be recognized by the speech recognition end.
[0015] According to another aspect of the embodiments of the present application, another voice separation device is provided. The device can include: a seventh obtaining unit configured to obtain a sequence of voice information by invoking a first interface, wherein the first interface comprises a first parameter, a parameter value of the first parameter is the sequence of voice information, the sequence of voice information comprises at least one voice information to be separated, and different voice information is from different utterance objects; a fourth processing unit configured to extract voice features of different utterance objects from the sequence of voice information to obtain a sequence of voice features; a fifth processing unit configured to perform gating processing on the voice features in the sequence of voice features according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local voice information and global voice information of different utterance objects, and an information granularity of the local voice information is smaller than an information granularity of the global voice information; an eighth obtaining unit configured to obtain voice mask information of different utterance objects based on the gating processing result, wherein the voice mask information is used to represent utterance attributes of the utterance objects; a fifth separation unit configured to separate voice information output by different utterance objects from the sequence of voice information based on the voice mask information of different utterance objects and the sequence of voice features; and an output unit configured to output the voice information output by different utterance objects by invoking a second interface, wherein the second interface comprises a second parameter, and a value of the second parameter is the voice information output by different utterance objects.
[0016] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided, which comprises a stored program, wherein the program, when executed, controls a device where the storage medium is located to perform any of the voice separation methods.
[0017] According to another aspect of the embodiments of the present application, a processor is also provided, which is configured to execute a program, wherein the program, when executed, performs any of the voice separation methods.
[0018] In the embodiment of the present application, a speech information sequence is acquired, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different pronunciation objects; speech features of different pronunciation objects are extracted from the speech information sequence to obtain a speech feature sequence; speech features in the speech feature sequence are processed by a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information; based on the gating processing result, speech mask information of different pronunciation objects is acquired, wherein the speech mask information is used to represent the pronunciation attribute of the pronunciation object; and based on the speech mask information of different pronunciation objects and the speech feature sequence, speech information output by different pronunciation objects is separated from the speech information sequence. That is, the embodiment of the present application processes speech features in the acquired speech information sequence by a local attention mechanism and a global attention mechanism to obtain local speech information and global speech information of different pronunciation objects, and based on the gating processing, the requirements for the local attention mechanism and the global attention mechanism are greatly reduced, so that not only global information can be directly processed, but also smaller local features can be processed, thereby realizing the technical effect that speech can be separated, and further solving the technical problem that speech cannot be separated. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:
[0020] Figure 1 FIG. 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a speech separation method according to an embodiment of the present application;
[0021] Figure 2 FIG. 2 is a structure block diagram of a computing environment according to an embodiment of the present application;
[0022] Figure 3 FIG. 3 is a structure block diagram of a service mesh according to an embodiment of the present application;
[0023] Figure 4 FIG. 4 is a flowchart of a speech separation method according to an embodiment of the present application;
[0024] Figure 5 FIG. 5 is a flowchart of another speech separation method according to an embodiment of the present application;
[0025] Figure 6 FIG. 6 is a flowchart of another speech separation method according to an embodiment of the present application;
[0026] Figure 7 is a flowchart of another voice separation method according to an embodiment of the present application;
[0027] Figure 8 is a flowchart of another voice separation method according to an embodiment of the present application;
[0028] Figure 9 is a schematic diagram of access of a computer device to a private network according to an embodiment of the present application;
[0029] Figure 10 is a schematic diagram of a deep network model based on an attention mechanism according to an embodiment of the present application;
[0030] Figure 11 is a schematic diagram of a local and global hybrid attention mechanism architecture based on a gating mechanism according to an embodiment of the present application;
[0031] Figure 12 is a schematic diagram of a convolution module according to an embodiment of the present application;
[0032] Figure 13 is a schematic diagram of a voice separation device according to an embodiment of the present application;
[0033] Figure 14 is a schematic diagram of another voice separation device according to an embodiment of the present application;
[0034] Figure 15 is a schematic diagram of another voice separation device according to an embodiment of the present application;
[0035] Figure 16 is a schematic diagram of another voice separation device according to an embodiment of the present application;
[0036] Figure 17 is a schematic diagram of another voice separation device according to an embodiment of the present application;
[0037] Figure 18 is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present application.
[0039] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0040] First, some of the nouns or terms that appear in the description of the embodiments of the application are applicable to the following explanations:
[0041] Speech separation can be to separate the mixed speech of multiple speakers and obtain the individual speech of all speakers;
[0042] Self-attention mechanism can be a sequence processing module algorithm used in deep learning models (such as Transformer models);
[0043] Deep learning algorithm can be a model modeling method based on multi-layer neural network;
[0044] Convolution can be a mathematical operator that generates a third function by two functions, which can represent the area of the curved trapezoid surrounded by the product function after flipping and translation;
[0045] Cocktail problem can refer to the problem that when multiple people communicate at the same time, if speech separation is not performed after being collected by the microphone, it will directly affect the speech recognition system or the degree of auditory perception and understanding.
[0046] Embodiment 1
[0047] According to the embodiments of the application, a speech separation method is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that herein.
[0048] The method embodiment provided by the embodiment of the application can be executed in a mobile terminal, a computer terminal or a similar operation device. Figure 1is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a speech separation method according to an embodiment of the present application. As shown in Figure 1 The computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .
[0049] It should be noted that the one or more processors 102 and / or other speech separation circuits described above can be referred to herein generally as "speech separation circuits". The speech separation circuits can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the speech separation circuits can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As referred to in the embodiments of the present application, the speech separation circuits serve as a processor control (for example, selection of a variable resistance terminal path connected to an interface).
[0050] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the speech separation method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the speech separation method described above. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory remotely disposed with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0051] The transport 106 is configured to receive or transmit data via a network. Examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transport 106 includes a network interface controller (NIC) that can connect to other network devices through a base station to communicate with the Internet. In one example, the transport 106 can be a radio frequency (RF) module that is configured to communicate with the Internet via wireless means.
[0052] The display can be a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0053] Figure 1 The illustrated hardware architecture diagram can be used as an exemplary diagram of the computer terminal 10 (or mobile device) described above, and can also be used as an exemplary diagram of the server described above, in an alternative embodiment, Figure 2 The above-described computer terminal 10 (or mobile device) is illustrated in a block diagram as an embodiment of a computing node in a computing environment 201. Figure 1 The illustrated computer terminal 10 (or mobile device) is illustrated in a block diagram as an embodiment of a computing node in a computing environment 201. Figure 2 The illustrated computer terminal 10 (or mobile device) is illustrated in a block diagram as an embodiment of a computing node in a computing environment 201. Figure 2 As illustrated, the computing environment 201 includes a plurality of computing nodes (e.g., servers) 210-1, 210-2,... running on a distributed network. Each computing node includes local processing and memory resources, and end users 202 can remotely run applications or store data in the computing environment 201. Applications can be provided as a plurality of services 220-1, 220-2, 220-3, and 220-4 in the computing environment 201, representing services "A", "D", "E", and "H", respectively.
[0054] The end users 202 can provide and access the services through a web browser or other software application on a client, and in some embodiments, the provisioning and / or requests of the end users 202 can be provided to an entry gateway 230. The entry gateway 230 can include a corresponding proxy to handle the provisioning and / or requests for the services (one or more of the services provided in the computing environment 201).
[0055] Services are provided or deployed in accordance with various virtualization technologies supported by the computing environment 201. In some embodiments, services can be provided in accordance with virtualization based on virtual machines (VMs), virtualization based on containers, and / or the like. Virtualization based on virtual machines can be emulating a real computer by initializing a virtual machine to execute programs and applications without directly accessing any actual hardware resources. While the virtual machine is virtualized, in accordance with virtualization based on containers, a container can be launched to virtualize an entire operating system (OS) so that multiple workloads can run on a single operating system instance.
[0056] In one embodiment of container-based virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, as shown in Figure 2 Service 220-2 can be equipped with one or more Pods 240-1, 240-2, …, 240-N (collectively, Pods) as shown. Each Pod can include a proxy 245 and one or more containers 242-1, 242-2, …, 242-M (collectively, containers). The one or more containers in a Pod handle requests related to one or more respective functions of the service, and the proxy 245 generally controls network functions related to the service, such as routing, load balancing, and the like. Other services can also be equipped with Pods similar to the Pods.
[0057] In operation, executing a user request from the end user 202 can require invoking one or more services in the computing environment 201, executing one or more functions of a service can require invoking one or more functions of another service. As shown in Figure 2 Service “A” 220-1 receives a user request from the end user 202 from the ingress gateway 230, service “A” 220-1 can invoke service “D” 220-2, which can request service “E” 220-3 to execute one or more functions.
[0058] The computing environment described above can be a cloud computing environment, in which allocation of resources is managed by a cloud service provider, allowing development of functionality without considering implementation, tuning, or scaling servers. The computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically scale independently, rather than scaling a single hardware device to handle potential loads.
[0059] In another alternative embodiment, Figure 3 A block diagram is shown to illustrate use of the above-described Figure 1 The computer terminal 10 (or mobile device) is shown as an embodiment of a service mesh.Figure 3 is a structural diagram of a service mesh according to an embodiment of the present application, as Figure 3 shown, the service mesh 300 is mainly used to facilitate secure and reliable communication between multiple microservices, which refers to decomposing an application into multiple smaller services or instances and running them on different clusters / machines.
[0060] As Figure 3 shown, the microservices can include an application service instance A and an application service instance B, which form a functional application layer of the service mesh 300. In an implementation, the application service instance A runs in the form of a container / process 308 on a machine / workload container group 314 (Pod), and the application service instance B runs in the form of a container / process 310 on a machine / workload container group 316 (Pod).
[0061] In an implementation, the application service instance A can be a commodity query service, and the application service instance B can be a commodity ordering service.
[0062] As Figure 3 shown, the application service instance A and a mesh proxy (sidecar) 303 coexist in the machine workload container group 614, and the application service instance B and a mesh proxy 305 coexist in the machine workload container 314. The mesh proxy 303 and the mesh proxy 305 form a data plane layer of the service mesh 300. Among them, the mesh proxy 303 and the mesh proxy 305 run in the form of a container / process 304 and a container / process 306, respectively, can receive a request 312 for commodity query service, and the mesh proxy 303 and the application service instance A can communicate bidirectionally, the mesh proxy 305 and the application service instance B can communicate bidirectionally. In addition, the mesh proxy 303 and the mesh proxy 305 can also communicate bidirectionally.
[0063] In an implementation, all traffic of the application service instance A is routed to the appropriate destination through the mesh proxy 303, and all network traffic of the application service instance B is routed to the appropriate destination through the mesh proxy 305. It should be noted that the network traffic mentioned herein includes but is not limited to Hyper Text Transfer Protocol (HTTP), Representational State Transfer (REST), G Remote Procedure Call (GRPC), and Redis, etc.
[0064] In an implementation, the function of the extended data plane layer can be implemented by writing a custom filter for the proxies (Envoy) in the service mesh 300. The service mesh proxy configuration can be to make the service mesh correctly proxy service traffic, implement service interworking and service governance. The mesh proxy 303 and the mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0065] As shown in Figure 3 , the service mesh 300 further includes a control plane layer. The control plane layer can be a group of services running in a dedicated namespace, hosted by the hosting control plane component 301 in the machine / Pod 302. As shown in Figure 3 , the hosting control plane component 301 communicates with the mesh proxy 303 and the mesh proxy 305 in a bidirectional manner. The hosting control plane component 301 is configured to perform some control management functions. For example, the hosting control plane component 301 receives the telemetry data transmitted by the mesh proxy 303 and the mesh proxy 305, and can further aggregate the telemetry data. The services, the hosting control plane component 301 can also provide a user-oriented application programming interface (API) to facilitate manipulation of network behavior, and provide configuration data to the mesh proxy 303 and the mesh proxy 305, etc.
[0066] In the above running environment, the present application provides a voice separation method as shown in Figure 4 . Figure 4 is a flow chart of a voice separation method according to an embodiment of the present application. As shown in Figure 4 , the method can include the following steps:
[0067] Step S402, obtaining a voice information sequence, wherein the voice information sequence includes at least one voice information to be subjected to voice separation, and different voice information comes from different sound producing objects.
[0068] In the technical solution provided in the step S402 of the present application, the speech information sequence can be obtained, wherein the speech information sequence (sequence X) can be mixed sound wave information including at least one speech information to be separated, and different speech information can come from different pronunciation objects. The pronunciation object can be a speaking object (speaker).
[0069] Optionally, speech information from different pronunciation objects can be obtained to obtain the speech information sequence.
[0070] In the step S404, the speech features of different pronunciation objects can be extracted from the speech information sequence to obtain the speech feature sequence.
[0071] In the technical solution provided in the step S404 of the present application, the speech features of different pronunciation objects can be extracted from the speech information sequence to obtain the speech feature sequence based on the speech features in the speech information sequence of different pronunciation objects. The speech features can be feature vectors and can be used to represent the content in the speech information sequence.
[0072] Optionally, the speech information of at least one pronunciation object can be obtained to obtain the speech information sequence, and the features of the speech information in the speech information sequence can be extracted by an encoder to obtain the speech features of different pronunciation objects, thereby obtaining the speech feature sequence.
[0073] In the step S406, the speech features in the speech feature sequence can be processed by the local attention mechanism and the global attention mechanism to obtain the gating processing result, wherein the gating processing result includes the local speech information and the global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information.
[0074] In the technical solution provided in the step S406 of the present application, the speech features in the speech feature sequence can be processed by the local attention mechanism and the global attention mechanism to obtain the gating processing result, wherein the gating processing result includes the local speech information and the global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information. The gating processing can include addition processing, multiplication processing and other processing modes of the speech features.
[0075] In the embodiment of the present application, a hybrid attention mechanism is proposed, which includes a local attention mechanism and a global attention mechanism. The relationship between the local features and the global features is learned by using the global attention mechanism and the local attention mechanism in the gating processing, thereby realizing the technical effect that the speech can be separated, and further solving the technical problem that the speech cannot be separated.
[0076] In step S408, speech mask information of different speakers is obtained based on the gating processing result, where the speech mask information is used to represent the speech attribute of the speaker.
[0077] In the technical solution provided in step S408, the speech mask information of different speakers can be obtained based on the gating processing result, where the speech mask information (individual speaker’s mask) can be used to represent the speech attribute of the speaker, and can be a mask matrix, for example, a time-frequency point mask matrix. It should be noted that this is only an example and the mask is not limited in particular.
[0078] In step S410, speech information output by different speakers is separated from the speech information sequence based on the speech mask information of different speakers and the speech feature sequence.
[0079] In the technical solution provided in step S410, speech information output by different speakers can be separated from the speech information sequence based on the speech mask information of different speakers and the speech feature sequence, for example, by multiplying the speech mask information and the speech feature sequence.
[0080] In the technical solution provided in steps S402-S410, the speech information sequence is obtained, where the speech information sequence includes at least one speech information to be separated, and different speech information comes from different speakers. The speech features of different speakers are extracted from the speech information sequence to obtain a speech feature sequence. The speech features in the speech feature sequence are gated according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, where the gating processing result includes local speech information and global speech information of different speakers, and the information granularity of the local speech information is smaller than that of the global speech information. The speech mask information of different speakers is obtained based on the gating processing result, where the speech mask information is used to represent the speech attribute of the speaker. The speech mask information of different speakers and the speech feature sequence are used to separate the speech information output by different speakers from the speech information sequence. In other words, the speech features in the obtained speech information sequence are gated according to the local attention mechanism and the global attention mechanism, and the local speech information and the global speech information of different speakers can be obtained. Based on the gating processing, the requirements for the local attention mechanism and the global attention mechanism are greatly reduced, so that not only the global information can be directly processed, but also the smaller local features can be processed, thereby achieving the technical effect of separating the speech, and solving the technical problem that the speech cannot be separated.
[0081] The above method of the embodiment is further described below.
[0082] As an optional implementation, the method comprises: the local attention mechanism comprises a single-head attention mechanism, the global attention mechanism comprises a linear attention mechanism, and the speech feature in the speech feature sequence is processed according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, comprising: converting the speech feature in the speech feature sequence according to the single-head attention mechanism to obtain local speech information; converting the speech feature in the speech feature sequence according to the linear attention mechanism to obtain global speech information; and performing gating processing on the local speech information and the global speech information to obtain the gating processing result.
[0083] In this embodiment, the local attention mechanism can comprise a single-head attention mechanism, and the global attention mechanism can comprise a linear attention mechanism. The single-head attention mechanism can be a self-attention mechanism (Self Attention), and the linear attention mechanism can be a simplified linear attention mechanism. The speech feature in the speech feature sequence can be converted according to the single-head attention mechanism to obtain local speech information, the speech feature in the speech feature sequence can be converted according to the linear attention mechanism to obtain global speech information, and the local speech information and the global speech information can be processed by gating to obtain a gating processing result.
[0084] In the embodiment of the application, by using gating, the multi-head attention mechanism (Multi-Head Attention) in the related art is simplified to a single-head attention mechanism, the speech feature in the speech feature sequence can be converted according to the single-head attention mechanism to obtain only local speech information, thereby achieving the purpose of reducing the amount of calculation. At the same time, the speech feature in the speech feature sequence can be converted according to the linear attention mechanism to obtain global information, thereby achieving the purpose of greatly simplifying the algorithm complexity.
[0085] As an optional implementation, the speech feature in the speech feature sequence is processed by convolution to obtain a speech feature matrix of a target dimension, and the speech feature in the speech feature sequence is converted according to the linear attention mechanism to obtain global speech information, comprising: converting the speech feature matrix according to the linear attention mechanism to obtain the global speech information.
[0086] In this embodiment, the speech feature in the speech feature sequence can be processed by convolution to obtain a speech feature matrix of a target dimension, and the speech feature in the speech feature sequence can be converted according to the linear attention mechanism to obtain global speech information.
[0087] Optionally, the speech features in the speech feature sequence can be processed by parallel convolution (Convolution Module), and a speech feature matrix of a target dimension can be obtained, for example, a speech feature matrix of S*A dimension (U and V), which can be determined by the following formula:
[0088] U = ConvM(X")
[0089] V = ConvM(X")
[0090] where ConvM can be used to represent the convolution module. The speech features can be converted into a speech feature matrix of a target dimension through a convolution layer, the speech features in the speech feature sequence can be converted according to a linear attention mechanism to obtain global speech information, and the global speech information of the speech feature matrix V and the speech feature matrix U can be determined by the following linearization form: global’ and U global’ ):
[0091] V global’ = Q'(βK' T V), U global’ = Q'(βK' T U)
[0092] where β can be a time scaling coefficient; Q' can be a corresponding feature sequence; K' T can be a key corresponding to the feature sequence.
[0093] As an optional implementation, the speech features in the speech feature sequence are converted according to a single-head attention mechanism to obtain local speech information, including: converting the block speech feature matrix of the speech feature matrix according to the single-head attention mechanism to obtain the local speech information.
[0094] In this embodiment, the speech feature matrix in the speech information sequence feature can be blocked, and the block speech feature matrix obtained after blocking can be converted according to a single-head attention mechanism to obtain local speech information. The block speech feature matrix can be a non-overlapping block speech feature matrix.
[0095] Optionally, the speech feature matrix can be divided into non-overlapping blocks of the same size using zero padding, and the split non-overlapping blocks can be converted according to a single-head attention mechanism to obtain local speech information (V local,h ' and U local,h '), which can be determined by the following formula:
[0096]
[0097] wherein, γ can be a scaling coefficient; RELU 2 can be a square of a rectified linear coefficient; can be calculated only once, V h and U h can be a block speech feature matrix obtained after block processing of the speech feature matrix.
[0098] In this embodiment, a square rectified linear coefficient (RELU 2 ) is used instead of a normalization exponential function (Softmax) in the multi-head attention mechanism, so as to further optimize the model performance.
[0099] As an optional implementation, the local speech information and the global speech information are subjected to gate processing to obtain a gate processing result, including: obtaining merged speech information between the global speech information and the local speech information; and performing gate processing on the merged speech information, the speech feature matrix and the speech feature sequence to obtain the gate processing result.
[0100] In this embodiment, the merged speech information between the global speech information and the local speech information can be obtained. The gate processing can be performed on the merged speech information, the speech feature matrix and the speech feature sequence to obtain the gate processing result.
[0101] Optionally, the merged speech information (V' and U') between the global speech information and the local speech information can be determined by the following formula:
[0102] V' = V global ' + V local ', U' = U global ' + U local '
[0103] Optionally, the gate processing can be performed on the merged speech information, the speech feature matrix and the speech feature sequence to obtain the gate processing result, wherein the gate processing can include: feature (element) activation processing feature summation processing and feature multiplication processing The gate processing result (O', O'', O) can be determined by the following formula:
[0104] wherein, V' = V * A
[0105] wherein, U' = A * U
[0106]
[0107] wherein, V' can be the merged speech information; A can be a convolution coefficient; U' can be the merged speech information; It can be used to activate functions for elements.
[0108] In this embodiment, gating processing can significantly reduce the requirements for attention mechanisms, thereby simplifying multi-head attention mechanisms into single-head attention mechanisms, and thus also significantly reducing the requirements for local and global attention mechanisms.
[0109] Optionally, for long sentences, the data processing process takes a long time. Therefore, in this embodiment of the invention, gating processing is used to combine the local speech feature matrix (U) and the global speech feature matrix (V) in an efficient and effective manner, thereby improving the model's efficiency in data processing.
[0110] As an optional implementation, the speech feature sequence is convolved to obtain a speech feature matrix of the target dimension, including: performing multiple convolution processes on the speech feature sequence to obtain speech feature matrices of different target dimensions.
[0111] In this embodiment, the speech feature sequence can be subjected to multiple convolution processes to obtain speech feature matrices of different target dimensions.
[0112] For example, a speech feature matrix with a target dimension of N*S can be obtained by performing pointwise convolution on a speech information sequence. Then, another pointwise convolution can be performed on the speech feature matrix with a target dimension of N*S to obtain a speech feature matrix with a target dimension of C*N*S.
[0113] As an optional implementation, the method further includes: normalizing the speech feature sequence to obtain a normalized speech result; encoding the normalized speech result to obtain a speech encoding result; performing convolution processing on the speech encoding result, and transforming the obtained convolution result to obtain a speech feature matrix of the original dimension; wherein, performing convolution processing on the speech features in the speech feature sequence to obtain a speech feature matrix of the target dimension includes: performing convolution processing on the speech feature matrix of the original dimension to obtain a speech feature matrix of the target dimension.
[0114] In this embodiment, the speech feature sequence can be normalized to obtain a normalized speech result, the normalized speech result can be encoded to obtain a speech encoding result, the speech encoding result can be convolved, the obtained convolution result can be transformed to obtain a speech feature matrix of the original dimension, and the speech feature matrix of the original dimension can be convolved to obtain a speech feature matrix of the target dimension.
[0115] Optionally, the speech information sequence output by the encoder can first pass through a linear layer for normalization processing (LayerNorm) to obtain a normalized speech result, where the normalized speech result can be a speech feature matrix. Positional encodings can be added to the normalized speech to obtain a speech encoding result, where the added positional encodings can be sinusoidal positional encodings (Sinusoidal Positional Encodings). This is only an example and is not limited. The speech encoding result with added positional encodings can be convoluted by pointwise convolution (Pointwise Convolution), and the obtained convolution result can be passed and reshaped (Reshape) to obtain a speech feature matrix of the original dimension. The speech feature matrix of the original dimension can be convoluted to obtain a speech feature matrix of a target dimension.
[0116] As an optional implementation, speech features of different pronunciation objects are extracted from the speech information sequence to obtain a speech feature sequence, including: convoluting the speech information sequence to obtain speech features of different pronunciation objects; and linearly processing the speech features of different pronunciation objects to obtain the speech feature sequence.
[0117] In this embodiment, the speech information sequence can be convoluted to obtain speech features of different pronunciation objects, where the speech features can be used to represent pronunciation attributes of the pronunciation objects. The speech features of different pronunciation objects can be linearly processed to obtain the speech feature sequence.
[0118] Optionally, the encoder can be composed of a one-dimensional (1 Dimension, 1D for short) convolution layer (Convolution) and a rectified linear unit (Rectified Linear Unit, ReLU for short). The rectified linear unit can be used to constrain the output speech feature sequence to be a non-negative value.
[0119] Optionally, it can be assumed that the kernel size of the encoder is K1, the step size is K1 / 2, and the number of filters in the encoder can be N. The input speech information sequence (X) to the encoder can be determined by the following formula to output the speech feature sequence (X’):
[0120] X’ = RELU (Conv 1D (X))
[0121] Where, RELU can be a linear processing parameter; and Conv 1D can be a one-dimensional convolution layer parameter.
[0122] As an optional implementation, based on the gating processing result, the speech mask information of different pronunciation objects is obtained, comprising: performing linear processing on the gating processing result, and performing convolution processing on the obtained linear processing result to obtain the speech mask information of different pronunciation objects.
[0123] In this embodiment, the gating processing result can be obtained, linear processing can be performed on the gating processing result, and convolution processing can be performed on the linear processing result to obtain the speech mask information of different pronunciation objects.
[0124] Optionally, the gating processing result can be obtained, rectified linear processing can be performed on the gating processing result, and point-by-point convolution can be performed on the linear processing result to obtain the speech mask information of different pronunciation objects.
[0125] As an optional implementation, based on the speech mask information of different pronunciation objects and the speech feature sequence, the speech information output by different pronunciation objects is separated from the speech information sequence, comprising: obtaining the product result between the speech mask information of different pronunciation objects and the speech feature sequence; and determining the product result as the speech information output by different pronunciation objects.
[0126] In the embodiment of the application, the speech mask information of different pronunciation objects is obtained, the product result between the speech mask information of different pronunciation objects and the speech feature sequence is calculated, and the product result can be determined as the speech information output by different pronunciation objects.
[0127] Optionally, the speech mask information (M i ) of different pronunciation objects and the speech feature sequence (X') are obtained, the product result of the speech mask information and the speech feature sequence (X i ") is determined, the product result can be determined as the speech information (X i ") output by different pronunciation objects, and the speech information output by different pronunciation objects can be determined by the following formula:
[0128] ”’
[0129] X i =M i *X
[0130] In the embodiment of the present application, the speech features in the obtained speech information sequence are subjected to gating processing according to the local attention mechanism and the global attention mechanism, so that the local speech information and the global speech information including different pronunciation objects can be obtained. Based on the gating processing, the requirements for the local attention mechanism and the global attention mechanism are greatly reduced, so that not only the global information can be directly processed, but also smaller local features can be processed, thereby realizing the technical effect that the speech can be separated, and thereby solving the technical problem that the speech cannot be separated.
[0131] The speech separation method will be further introduced below from the scenario of using the speech separation model.
[0132] Figure 5 is a flowchart of another speech separation method according to an embodiment of the present application. As shown in Figure 5 , the method can include the following steps:
[0133] In step S502, a speech information sequence is obtained, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different pronunciation objects.
[0134] In step S504, a speech separation model is called, wherein the speech separation model is obtained based on training of the local attention mechanism and the global attention mechanism.
[0135] In the technical solution provided in the above step S504 of the present application, the speech information sequence is obtained, so that the speech separation model can process the speech information sequence, and the separation of the speech in the speech information sequence is completed. The speech separation model can be a model obtained based on training of the local attention mechanism and the global attention mechanism, such as a deep neural network model obtained based on training of the local attention mechanism and the global attention mechanism.
[0136] Optionally, the speech separation model can be a deep neural network model including an encoder, a decoder and a masker, can be a model trained based on a hybrid attention mechanism of the local attention mechanism and the global attention mechanism, and can be used for separating the speech information in the mixed speech information.
[0137] The embodiment of the present application proposes a deep network model algorithm based on an attention mechanism, which can be based on a model framework of a gated attention mechanism and modeling of local data features, trained based on the local attention mechanism and the global attention mechanism to obtain a speech separation model. The speech separation model is trained based on the local attention mechanism and the global attention mechanism, which not only simplifies the complexity of the algorithm, but also directly processes the global information and processes smaller local features, improves the effect of separating the speech, and thereby better solves the problem of speech separation.
[0138] In step S506, the speech feature sequence is obtained by using the speech separation model to extract speech features of different pronunciation objects from the speech information sequence, and the speech features in the speech feature sequence are processed according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information.
[0139] In the technical solution provided in step S506, the speech separation model can be used to process the speech information sequence, extract speech of different pronunciation objects from the speech information sequence, extract features from the speech information sequence, and extract speech features of different pronunciation objects from the speech information sequence. Based on the speech features in the speech information sequence of different pronunciation objects, a speech feature sequence is obtained. The speech features in the speech feature sequence can be processed according to the local attention mechanism and the global attention mechanism, respectively, to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information. The gating processing can include addition processing, multiplication processing and other processing modes of the speech features.
[0140] Optionally, the speech information of at least one pronunciation object is obtained to obtain a speech information sequence, the features of the speech information in the speech information sequence can be extracted by an encoder in the speech separation model to obtain speech features of different pronunciation objects, thereby obtaining a speech feature sequence. The speech features in the speech feature sequence can be processed according to the local attention mechanism and the global attention mechanism to obtain a gating processing result.
[0141] In step S508, the speech mask information of different pronunciation objects is obtained based on the gating processing result, wherein the speech mask information is used to represent the pronunciation attribute of the pronunciation object.
[0142] In step S510, the speech information output by different pronunciation objects is separated from the speech information sequence based on the speech mask information of different pronunciation objects and the speech feature sequence.
[0143] By the steps S502 to S510, the speech information sequence is obtained, wherein the speech information sequence comprises at least one speech information to be separated, and different speech information comes from different pronunciation objects; the speech separation model is called, wherein the speech separation model is obtained by training based on the local attention mechanism and the global attention mechanism; the speech features of the different pronunciation objects are extracted from the speech information sequence by using the speech separation model, and the speech feature sequence is obtained, and the speech features in the speech feature sequence are processed by the local attention mechanism and the global attention mechanism, and the processing result is obtained, wherein the processing result comprises local speech information and global speech information of the different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information; the speech mask information of the different pronunciation objects is obtained based on the processing result, wherein the speech mask information is used to represent the pronunciation attribute of the pronunciation object; and the speech information output by the different pronunciation objects is separated from the speech information sequence based on the speech mask information of the different pronunciation objects and the speech feature sequence, thereby realizing the technical effect that the speech cannot be separated, and solving the technical problem that the speech cannot be separated.
[0144] The speech separation method is further introduced from the speech playback scene.
[0145] Figure 6 is a flowchart of another speech separation method according to an embodiment of the present application. As shown in Figure 6 , the method can include the following steps:
[0146] In step S602, the speech features of the different pronunciation objects are extracted from the obtained speech information sequence, and the speech feature sequence is obtained, wherein the speech information sequence comprises at least one speech information to be separated, and different speech information comes from different pronunciation objects.
[0147] In step S604, the speech features in the speech feature sequence are processed by the local attention mechanism and the global attention mechanism, and the processing result is obtained, wherein the processing result comprises local speech information and global speech information of the different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information.
[0148] In step S606, the speech mask information of the different pronunciation objects is obtained based on the processing result, wherein the speech mask information is used to represent the pronunciation attribute of the pronunciation object.
[0149] In step S608, the speech information output by the different pronunciation objects is separated from the speech information sequence based on the speech mask information of the different pronunciation objects and the speech feature sequence.
[0150] Step S610, respectively playing the speech information output by different pronunciation objects.
[0151] By the steps S602 to S610 of the present application, the speech features of different pronunciation objects are extracted from the obtained speech information sequence, to obtain a speech feature sequence, wherein the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different pronunciation objects; the speech features in the speech feature sequence are subjected to gating processing according to a local attention mechanism and a global attention mechanism, to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information; based on the gating processing result, speech mask information of different pronunciation objects is obtained, wherein the speech mask information is used to represent the pronunciation attribute of the pronunciation object; based on the speech mask information of different pronunciation objects and the speech feature sequence, the speech information output by different pronunciation objects is separated from the speech information sequence; and the speech information output by different pronunciation objects is respectively played, thereby achieving the technical effect that speech cannot be subjected to speech separation, and solving the technical problem that speech cannot be subjected to speech separation.
[0152] The speech separation method will be further introduced from the speech recognition scene.
[0153] Figure 7 is a flowchart of another speech separation method according to an embodiment of the present application. As shown in Figure 7 , the method can include the following steps:
[0154] Step S702, from the obtained speech information sequence, the speech features of different pronunciation objects are extracted, to obtain a speech feature sequence, wherein the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different pronunciation objects.
[0155] Step S704, the speech features in the speech feature sequence are subjected to gating processing according to a local attention mechanism and a global attention mechanism, to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information.
[0156] Step S706, based on the gating processing result, speech mask information of different pronunciation objects is obtained, wherein the speech mask information is used to represent the pronunciation attribute of the pronunciation object.
[0157] Step S708, based on the speech mask information of different pronunciation objects and the speech feature sequence, the speech information output by different pronunciation objects is separated from the speech information sequence.
[0158] In step S710, the speech information output by the different pronunciation objects is input to the speech recognition end, and the speech information is used for recognition by the speech recognition end.
[0159] In the technical solution provided in step S710, the speech information output by the different pronunciation objects is input to the speech recognition end, and the speech information is used for recognition by the speech recognition end, and the speech recognition end can respond to the recognition result.
[0160] For example, the speech recognition end can be an intelligent voice assistant. When the speech recognition end recognizes the speech information output by the different pronunciation objects, the speech information of the owner can be recognized, and the content of the speech information of the owner can be recognized and the corresponding action can be performed. For example, the speech information of the owner is "turn on the music player", and when the speech recognition end recognizes the speech information of the owner, the instruction to turn on the music player can be executed.
[0161] It should be noted that the above scenario is only an example, and the speech recognition end is not specifically limited herein, and the use scenario of the speech separation method is not specifically limited herein. The scenario of speech separation should be within the protection scope of the embodiments of the present application.
[0162] In the above steps S702-S710, the speech features of the different pronunciation objects are extracted from the obtained speech information sequence to obtain a speech feature sequence, the speech information sequence includes at least one speech information to be separated, and the different speech information comes from different pronunciation objects. The speech features in the speech feature sequence are processed according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, the gating processing result includes local speech information and global speech information of the different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information. Based on the gating processing result, speech mask information of the different pronunciation objects is obtained, and the speech mask information is used to represent the pronunciation attribute of the pronunciation object. Based on the speech mask information of the different pronunciation objects and the speech feature sequence, the speech information output by the different pronunciation objects is separated from the speech information sequence. The speech information output by the different pronunciation objects is input to the speech recognition end, and the speech information is used for recognition by the speech recognition end. The technical effect of being unable to separate the speech is achieved, and the technical problem of being unable to separate the speech is solved.
[0163] The embodiments of the present application also provide another speech separation method, which can be applied to a software service side (Software-as-a-Service, SaaS for short).
[0164] Figure 8is a flowchart of another voice separation method according to an embodiment of the present application, as shown in Figure 8 The method can include the following steps.
[0165] In step S802, a voice information sequence is obtained by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter is the voice information sequence, the voice information sequence includes at least one voice information to be separated, and different voice information comes from different sound objects.
[0166] In the technical solution provided in the above step S802 of the present application, the first interface can be an interface for data interaction between a server and a user terminal. The user terminal can obtain the voice information sequence by calling the first interface. The voice information sequence is a first parameter of the first interface, which achieves the purpose of obtaining the voice information sequence. The voice information sequence can include at least one voice information to be separated, and different voice information can come from different sound objects.
[0167] In step S804, voice features of different sound objects are extracted from the voice information sequence to obtain a voice feature sequence.
[0168] In step S806, the voice features in the voice feature sequence are processed according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local voice information and global voice information of different sound objects, and the information granularity of the local voice information is smaller than that of the global voice information.
[0169] In step S808, based on the gating processing result, voice mask information of different sound objects is obtained, wherein the voice mask information is used to represent the sound properties of the sound objects.
[0170] In step S810, based on the voice mask information of different sound objects and the voice feature sequence, voice information output by different sound objects is separated from the voice information sequence.
[0171] In step S812, the voice information output by different sound objects is output by calling a second interface, wherein the second interface includes a second parameter, and the value of the second parameter is the voice information output by different sound objects.
[0172] In the technical solution provided in the above step S812 of the present application, the second interface can be an interface for data interaction between a server and a user terminal. The server can issue the voice information output by different sound objects to the client, so that the client can output the voice information output by different sound objects to the second interface as a parameter of the second interface, thereby achieving the purpose of issuing the voice information to the user terminal.
[0173] Figure 9 is a schematic diagram of access of a computer device to a private network according to an embodiment of the application, as Figure 9 As shown in the figure, the voice information sequence can be obtained by calling the first interface, and the computer device performs: step S902, extracting voice features of different pronunciation objects from the voice information sequence to obtain a voice feature sequence; step S904, performing gating processing on the voice features in the voice feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result including local voice information and global voice information of different pronunciation objects; step S906, obtaining voice mask information for representing pronunciation attributes of different pronunciation objects based on the gating processing result; and step S908, separating voice information output by different pronunciation objects from the voice information sequence based on the voice mask information of different pronunciation objects and the voice feature sequence; and the voice information output by different pronunciation objects can be output by calling the second interface.
[0174] Optionally, the platform can output the voice information output by different pronunciation objects by calling the second interface, where the second interface can be used to send the target domain name to the client, so that the client sends the voice information output by different pronunciation objects.
[0175] The embodiment of the application obtains the voice information sequence by calling the first interface, where the first interface includes a first parameter, the parameter value of the first parameter is the voice information sequence, the voice information sequence includes at least one voice information to be separated, and different voice information comes from different pronunciation objects; the voice features of different pronunciation objects are extracted from the voice information sequence to obtain a voice feature sequence; the voice features in the voice feature sequence are gated according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, where the gating processing result includes local voice information and global voice information of different pronunciation objects, and the information granularity of the local voice information is smaller than that of the global voice information; the voice mask information of different pronunciation objects is obtained based on the gating processing result, where the voice mask information is used to represent the pronunciation attributes of the pronunciation objects; the voice information output by different pronunciation objects is separated from the voice information sequence based on the voice mask information of different pronunciation objects and the voice feature sequence; and the voice information output by different pronunciation objects is output by calling the second interface, where the second interface includes a second parameter, the value of the second parameter is the voice information output by different pronunciation objects, thereby achieving the technical effect of separating voice, and further solving the technical problem of being unable to separate voice.
[0176] Embodiment 2
[0177] Speech separation can separate a single source speech from an overlapping mixed speech, and when multiple people communicate at the same time, if speech separation is not performed, it will directly affect the speech recognition system or the auditory perception and understanding. Therefore, in order to improve the recognition effect and auditory perception, the speech of multiple speakers mixed together is usually separated through speech separation to obtain a separation result, and the separation result can be used as an input signal of speech recognition or directly played to a listener.
[0178] In the related art, an end-to-end speech separation model (Wavesplit model) is proposed by speaker clustering, which uses the label of the speech content of an additional person in training, thereby increasing the training cost, and the method is only based on a convolutional network, and there is still a problem that the global information of the speech information sequence cannot be processed.
[0179] In another related art, a speech separation (SepFormer) model is proposed, which uses a multi-head attention mechanism, but the method only truncates a long sequence into a short sequence, and then performs intra-sequence and inter-sequence attention processing, and the global processing mode is only through implicit non-direct interaction, and there is still a problem that the global information of the speech information sequence cannot be processed, and the method also has the technical problem of low efficiency of speech separation for speech.
[0180] To solve the above problems, an embodiment of the present application proposes a deep network model algorithm based on an attention mechanism, which is based on a model framework of a gated attention mechanism and modeling of local data features, thereby not only simplifying the complexity of the algorithm, but also directly processing global information and processing smaller local features, improving the effect of speech separation for speech, thereby better solving the problem of speech separation.
[0181] The deep network model algorithm based on the attention mechanism proposed in the embodiment of the present application will be further introduced below.
[0182] Figure 10 is a schematic diagram of a deep network model based on an attention mechanism according to the embodiment of the present application, as Figure 10 shown, the deep network model based on the attention mechanism (MOSSFORMER model) can include an encoder (Encoder), a decoder (Decoder) and a masker (Masking Net). The encoder and the decoder can be used for feature extraction and waveform reconstruction in speech information, respectively. The masker is used to map the output of the encoder to a set of masks.
[0183] In this embodiment, as Figure 10As shown, a mixture sequence of speech information is obtained, and the mixture sequence of speech information is input into an encoder, where the encoder can be composed of a one-dimensional convolution layer and a rectified linear unit. The rectified linear unit can be used to constrain the output speech feature sequence to be a non-negative value.
[0184] Optionally, it can be assumed that the kernel size of the encoder is K1, the step size is K1 / 2, and the number of filters in the encoder can be N, then the input speech information sequence (X) into the encoder can be determined by the following formula to output the speech feature sequence (X’):
[0185] X’ = RELU (Conv 1D (X))
[0186] Optionally, the sequence X’ can be multiplied by the mask (M i ) of each speaker element by element, so that the separated feature sequence (X i ”) can be obtained, and the feature sequence (X i ”) can be determined by the following formula:
[0187] ”’
[0188] X i = M i *X
[0189] The separated feature sequence can finally be decoded by a one-dimensional transposed convolution layer (1D TransposedConvolution) in the decoder to obtain a speech information sequence of each speaker, where the speech information sequence of each speaker can be represented in a separated waveform (Separated Source), and the obtained separated waveform can be represented in the following way
[0190]
[0191] Optionally, the decoder can be a one-bit transposed convolution layer, and the decoder can use the same size of kernel and stride as the encoder.
[0192] In this embodiment, as shown in Figure 10 , a masker can be used for nonlinear mapping of the speech information sequence (X’) output by the encoder.
[0193] Optionally, as shown in Figure 10As shown, the speech information sequence output by the encoder can first pass through a linear layer for normalization processing to obtain a normalized speech result. The normalized speech result can be added with position encoding. The sequence added with the position encoding can be transmitted through pointwise convolution and reshaped, and then transmitted to a gated mechanism-based local and global hybrid attention mechanism architecture (MossFormer Block) for processing. The result processed by the gated mechanism-based local and global hybrid attention mechanism architecture can be output to a rectified linear unit for another pointwise convolution. The dimension of the obtained sequence RN*S can be expanded to RC*N*S. After the pointwise convolution and the gated linear unit (GLU), the sequence can pass through another pointwise convolution and the rectified linear unit to obtain a masked speech information sequence (M). There is a corresponding masked speech information sequence for each speaker object. Then, the masked speech information sequence corresponding to each speaker object is output to the decoder for processing.
[0194] In this embodiment, as shown in Figure 10 The input facilitates training. N gated mechanism-based local and global hybrid attention mechanism architectures can be set. The output of the current gated mechanism-based local and global hybrid attention mechanism architecture can be transmitted as input to the next gated mechanism-based local and global hybrid attention mechanism architecture, until the last gated mechanism-based local and global hybrid attention mechanism architecture outputs the processed data to the rectified linear unit.
[0195] In this embodiment, the speech information sequence can be processed by a convolution module and an attention gate mechanism. The convolution module can use linear projection and deep convolution processing. The attention gate mechanism can include local attention mechanism, global attention mechanism, and gating operation. The modeling capability of the gated mechanism-based local and global hybrid attention mechanism architecture is improved by the convolution module and the gating structure. The use of the gating structure effectively promotes the joint attention of the local and global.
[0196] Figure 11 FIG. 1 is a schematic diagram of a gated mechanism-based local and global hybrid attention mechanism architecture according to an embodiment of the present application. As shown in Figure 11 The gated mechanism-based local and global hybrid attention mechanism architecture can include a convolution module, a scale and offset rope module, a local and global joint attention module, and a gating operation module.
[0197] In the embodiment of the application, the dense layer in the Gated Attention Unit (GAU) can be replaced by a convolution module, thereby improving the extraction efficiency of fine-grained local features. Figure 12 is a schematic diagram of a convolution module according to an embodiment of the application, as Figure 12 shown, the convolution module can normalize and project the input speech information sequence through a linear layer, can perform linear processing on the normalized data through an activation layer (SiLU Activation), can perform feature convolution on the sequence through one-dimensional deep convolution, and can complete the training and regularization of the convolution module through random dropout processing (Dropout) on the data after feature convolution.
[0198] Optionally, the gating operation module can be a triple gating to enhance the model capability, and it should be noted that the number of "gates" in the gating module is not specifically limited. As Figure 11 shown, the input (X'') of the local and global hybrid attention mechanism architecture based on the gating mechanism can be obtained, and the convolution layer processing results (U and V) can be obtained after the convolution layer 1101 and the convolution layer 1102 are processed, respectively. The convolution layer processing results can be determined by the following formula:
[0199] U = ConvM (X'')
[0200]
[0201] wherein, ConvM can be used to represent the convolution module. Through the convolution layer, the speech feature can be converted from N dimensions to a 2N-dimensional speech feature matrix. The processing result of the convolution layer can be processed by the gating operation module to obtain the gating processing result (O', O'', O)
[0202]
[0203]
[0204]
[0205] In this embodiment, the gating processing can greatly reduce the requirement for the attention mechanism, thereby achieving the purpose of simplifying the multi-head attention mechanism into a single-head attention mechanism, and also greatly reducing the requirement for the local and global attention mechanisms.
[0206] Optionally, for long sentences, the data processing process requires a long time, and therefore, in the embodiment of the application, the local (U) and the whole (V) can be combined in an efficient and effective manner through the gating operation module, thereby improving the efficiency of the model in processing data.
[0207] In this embodiment, a hybrid attention mechanism architecture can be used, and in the local attention mechanism, only a single-head attention mechanism is used. A simplified linear attention mechanism can be used in the overall attention mechanism, and a single-head attention mechanism can be used in the local attention mechanism.
[0208] Optionally, as shown in Figure 11 , an input sentence (X") can be obtained first, and the X" is processed by the convolution layer 1103 to obtain a shared representation Z. The representation Z can be calculated by the following formula:
[0209] Z = ConvM (X")
[0210] As shown in Figure 11 , the Z output by the convolution layer can be obtained by the obtaining module in the offset & time scaling & obtaining module, and the Z can be shared to obtain the local and overall query words Q and the key K. In order to use the overall linear attention mechanism, the global speech information of the speech feature matrix V and the speech feature matrix U can be described by the following linearization form:
[0211] V global ’ = Q’ (βK T V), U global ’ = Q’ (βK T U)
[0212] Where β can be a time scaling coefficient.
[0213] Optionally, in order to calculate the local attention, the V, U, Q and K can be divided into H non-overlapping blocks of size P in a zero padding manner. The non-overlapping blocks can be converted according to the single-head attention mechanism to obtain the local speech information (V local,h ’ and U local,h ’), which can be determined by the following formula:
[0214]
[0215] Where γ can be a scaling coefficient.
[0216] In this embodiment, the square rectified linear coefficient (RELU 2 ) is used instead of the normalization exponential function (softmax) in the multi-head attention mechanism (Multi-HeadAttention), so that the model performance can be further optimized.
[0217] Optionally, the global speech information and the local speech information can be added together to form the final joint attention of V' and the sequence U':
[0218] V' = V global ’ + Vlocal ', U' = U global ' + U local '
[0219] In the embodiment of the present application, in order to better solve the attention mechanism modeling capability of the long sequence, a local and global hybrid attention mechanism architecture based on a gating mechanism is proposed. The gating mechanism can greatly reduce the requirement for the attention mechanism, and can simplify the multi-head attention mechanism into a single-head attention mechanism, thereby greatly reducing the requirement for the local and global attention mechanism. In the local attention mechanism, only a single-head attention mechanism can be used, thereby achieving the purpose of significantly reducing the calculation amount. At the same time, in the global attention mechanism, a simplified linear attention mechanism can be used to achieve the purpose, thereby greatly simplifying the complexity of the algorithm and directly processing the global information.
[0220] The attention mechanism mainly processes global information and does not process much on smaller local features, and cannot effectively extract the characteristics of short-term changes in speech. In order to make up for this deficiency, the embodiment of the present application further proposes a convolution processing module, which uses a deep convolution layer to extract local features, and by fusing the convolution processing module and the attention mechanism based on the gating mechanism, the technical effect of speech separation of speech is achieved, and the technical problem of being unable to perform speech separation on speech is solved.
[0221] It should be noted that, for each of the above method embodiments, in order to simply describe, each is described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0222] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method of each embodiment of the present application.
[0223] Embodiment 3
[0224] According to the embodiment of the present application, a method for implementing the aboveFigure 4 The speech separation device of the speech separation method.
[0225] Figure 13 is a schematic diagram of a speech separation device according to an embodiment of the present application. As shown, the speech separation device 1300 can include a first obtaining unit 1302, a first extracting unit 1304, a first processing unit 1306, a second obtaining unit 1308, and a first separating unit 1310. Figure 13
[0226] The first obtaining unit 1302 is configured to obtain a speech information sequence, where the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different utterance objects.
[0227] The first extracting unit 1304 is configured to extract speech features of different utterance objects from the speech information sequence to obtain a speech feature sequence.
[0228] The first processing unit 1306 is configured to perform gate processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gate processing result, where the gate processing result includes local speech information and global speech information of different utterance objects, and the information granularity of the local speech information is smaller than that of the global speech information.
[0229] The second obtaining unit 1308 is configured to obtain speech mask information of different utterance objects based on the gate processing result, where the speech mask information is used to represent utterance attributes of the utterance objects.
[0230] The first separating unit 1310 is configured to separate speech information output by different utterance objects from the speech information sequence based on the speech mask information of different utterance objects and the speech feature sequence.
[0231] It should be noted that the first obtaining unit 1302, the first extracting unit 1304, the first processing unit 1306, the second obtaining unit 1308, and the first separating unit 1310 correspond to steps S402 to S410 in Embodiment 1, and the five units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b, …, 102n), or the above units can be run in the computer terminal 10 provided in Embodiment 1 as a part of the device.
[0232] According to an embodiment of the present application, a computer program product is also provided, which includes a computer program and a computer readable storage medium. The computer program is configured to implement the above-mentioned Figure 5 The voice separation device of the voice separation method.
[0233] Figure 14 is a schematic diagram of another voice separation device according to an embodiment of the present application, as Figure 14 shown, the voice separation device 1400 can include a third acquisition unit 1402, a first calling unit 1404, a second extraction unit 1406, a fourth acquisition unit 1408, and a second separation unit 1410.
[0234] The third acquisition unit 1402 is configured to acquire a voice information sequence, wherein the voice information sequence includes at least one voice information to be subjected to voice separation, and different voice information comes from different sound production objects.
[0235] The first calling unit 1404 is configured to call a voice separation model, wherein the voice separation model is obtained based on local attention mechanism and global attention mechanism.
[0236] The second extraction unit 1406 is configured to extract voice features of different sound production objects from the voice information sequence using the voice separation model to obtain a voice feature sequence, and perform gate processing on the voice features in the voice feature sequence according to the local attention mechanism and the global attention mechanism to obtain a gate processing result, wherein the gate processing result includes local voice information and global voice information of different sound production objects, and the information granularity of the local voice information is smaller than that of the global voice information.
[0237] The fourth acquisition unit 1408 is configured to acquire voice mask information of different sound production objects based on the gate processing result, wherein the voice mask information is used to represent the sound production attributes of the sound production objects.
[0238] The second separation unit 1410 is configured to separate voice information output by different sound production objects from the voice information sequence based on the voice mask information of different sound production objects and the voice feature sequence.
[0239] It should be noted that the third acquisition unit 1402, the first calling unit 1404, the second extraction unit 1406, the fourth acquisition unit 1408, and the second separation unit 1410 correspond to steps S502 to S510 in Embodiment 1, and the five units have the same instances and application scenarios as the corresponding steps, but are not limited to the above-mentioned Embodiment 1. It should be noted that the above-mentioned units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b, …, 102n), and the above-mentioned units can also be a part of the device and can run in the computer terminal 10 provided in Embodiment 1.
[0240] According to the embodiment of the present application, a voice separation device for implementing the voice separation method shown in the above embodiment is further provided. Figure 6
[0241] Figure 15 is a schematic diagram of another voice separation device according to the embodiment of the present application, as shown in the figure, the voice separation device 1500 can include a third extraction unit 1502, a second processing unit 1504, a fifth acquisition unit 1506, a third separation unit 1508 and a playing unit 1510. Figure 15
[0242] The third extraction unit 1502 is configured to extract voice features of different utterance objects from the acquired voice information sequence to obtain a voice feature sequence, wherein the voice information sequence includes at least one voice information to be subjected to voice separation, and the different voice information comes from different utterance objects.
[0243] The second processing unit 1504 is configured to perform gate processing on the voice features in the voice feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gate processing result, wherein the gate processing result includes local voice information and global voice information of different utterance objects, and the information granularity of the local voice information is smaller than that of the global voice information.
[0244] The fifth acquisition unit 1506 is configured to acquire voice mask information of different utterance objects based on the gate processing result, wherein the voice mask information is used to represent the utterance attributes of the utterance objects.
[0245] The third separation unit 1508 is configured to separate voice information output by different utterance objects from the voice information sequence based on the voice mask information of different utterance objects and the voice feature sequence.
[0246] The playing unit 1510 is configured to play the voice information output by different utterance objects respectively.
[0247] It should be noted that the above-mentioned third extraction unit 1502, second processing unit 1504, fifth acquisition unit 1506, third separation unit 1508 and playing unit 1510 correspond to steps S602 to S610 in Embodiment 1, and the five units have the same instances and application scenarios as the corresponding steps, but are not limited to the contents disclosed in the above-mentioned embodiment 1. It should be noted that the above-mentioned units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b, …, 102n), and the above-mentioned units can also be a part of the device and can run in the computer terminal 10 provided in Embodiment 1.
[0248] According to an embodiment of the present application, a voice separation device for implementing the voice separation method shown in the above is also provided, which can be applied to the scenario of voice playback. Figure 7
[0249] Figure 16 is a schematic diagram of another voice separation device according to an embodiment of the present application. As shown in the figure, the voice separation device 1600 can include a fourth extraction unit 1602, a third processing unit 1604, a sixth acquisition unit 1606, a fourth separation unit 1608 and an input unit 1610. Figure 16
[0250] The fourth extraction unit 1602 is configured to extract voice features of different utterance objects from the acquired voice information sequence to obtain a voice feature sequence, wherein the voice information sequence includes at least one voice information to be subjected to voice separation, and the different voice information comes from different utterance objects.
[0251] The third processing unit 1604 is configured to perform gating processing on the voice features in the voice feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local voice information and global voice information of the different utterance objects, and the information granularity of the local voice information is smaller than that of the global voice information.
[0252] The sixth acquisition unit 1606 is configured to acquire voice mask information of the different utterance objects based on the gating processing result, wherein the voice mask information is used to represent the utterance attributes of the utterance objects.
[0253] The fourth separation unit 1608 is configured to separate the voice information output by the different utterance objects from the voice information sequence based on the voice mask information of the different utterance objects and the voice feature sequence.
[0254] The input unit 1610 is configured to input the voice information output by the different utterance objects to a voice recognition end, wherein the voice information is used to be recognized by the voice recognition end.
[0255] It should be noted that the fourth extraction unit 1602, the third processing unit 1604, the sixth acquisition unit 1606, the fourth separation unit 1608 and the input unit 1610 correspond to steps S702 to S710 in Embodiment 1, and the five units have the same instances and application scenarios as the corresponding steps, but are not limited to the above-mentioned embodiment 1. It should be noted that the above-mentioned units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b, …, 102n), or the above-mentioned units can be run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0256] According to an embodiment of the present application, a voice separation device for implementing the above-mentioned voice separation method is also provided. Figure 8 The voice separation device can be applied to a voice recognition scene.
[0257] Figure 17 is a schematic diagram of another voice separation device according to an embodiment of the present application. As shown in Figure 17 The voice separation device 1700 can include a seventh acquisition unit 1702, a fourth processing unit 1704, a fifth processing unit 1706, an eighth acquisition unit 1708, a fifth separation unit 1710 and an output unit 1712.
[0258] The seventh acquisition unit 1702 is configured to acquire a voice information sequence by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter is the voice information sequence, the voice information sequence includes at least one voice information to be subjected to voice separation, and different voice information comes from different pronunciation objects.
[0259] The fourth processing unit 1704 is configured to extract voice features of different pronunciation objects from the voice information sequence to obtain a voice feature sequence.
[0260] The fifth processing unit 1706 is configured to perform gating processing on the voice features in the voice feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local voice information and global voice information of different pronunciation objects, and the information granularity of the local voice information is smaller than that of the global voice information.
[0261] The eighth acquisition unit 1708 is configured to acquire voice mask information of different pronunciation objects based on the gating processing result, wherein the voice mask information is used to represent the pronunciation attribute of the pronunciation object.
[0262] The fifth separation unit 1710 is configured to separate speech information output by different speech objects from the speech information sequence based on the speech mask information and the speech feature sequence of the different speech objects.
[0263] The output unit 1712 is configured to output the speech information output by the different speech objects by calling a second interface, where the second interface includes a second parameter, and a value of the second parameter is the speech information output by the different speech objects.
[0264] It should be noted that the seventh acquisition unit 1702, the fourth processing unit 1704, the fifth processing unit 1706, the eighth acquisition unit 1708, the fifth separation unit 1710, and the output unit 1712 correspond to steps S802 to S812 in Embodiment 1, and the six units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b,..., 102n), or the units can be run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0265] In the speech separation device in this embodiment, the speech features in the obtained speech information sequence are subjected to the gating processing according to the local attention mechanism and the global attention mechanism, the local speech information and the global speech information including different speech objects can be obtained, and based on the gating processing, the requirements for the local attention mechanism and the global attention mechanism are greatly reduced, so that not only the global information can be directly processed, but also smaller local features can be processed, the technical effect that speech can be separated is achieved, and the technical problem that speech cannot be separated is solved.
[0266] Embodiment 4
[0267] Embodiments of the present application can provide a processor, which can include a computer terminal, and the computer terminal can be any one of computer terminal devices in a computer terminal group. Alternatively, in this embodiment, the computer terminal can be replaced by a mobile terminal or other terminal device.
[0268] Alternatively, in this embodiment, the computer terminal can be located in at least one network device of a plurality of network devices in a computer network.
[0269] In this embodiment, the computer terminal can execute program codes of the following steps in the speech separation method of the application: obtaining a speech information sequence, wherein the speech information sequence comprises at least one speech information to be subjected to speech separation, and different speech information comes from different sound production objects; extracting speech features of different sound production objects from the speech information sequence to obtain a speech feature sequence; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different sound production objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound production objects based on the gating processing result, wherein the speech mask information is used to represent sound production attributes of the sound production objects; and separating speech information output by different sound production objects from the speech information sequence based on the speech mask information of different sound production objects and the speech feature sequence.
[0270] Optionally, Figure 18 is a structural block diagram of a computer terminal according to an embodiment of the application. As shown in the figure, the computer terminal A can include one or more (only one is shown in the figure) processors 1802, a memory 1804, and a transmission device 1806. Figure 18
[0271] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the speech separation method and device in the embodiments of the application. The processor performs various functions and predictions by running the software programs and modules stored in the memory, that is, implements the speech separation method described above. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the computer terminal A through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0272] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining a speech information sequence, wherein the speech information sequence comprises at least one speech information to be subjected to speech separation, and different speech information comes from different pronunciation objects; extracting speech features of different pronunciation objects from the speech information sequence to obtain a speech feature sequence; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different pronunciation objects based on the gating processing result, wherein the speech mask information is used to represent the pronunciation attribute of the pronunciation object; and separating speech information output by different pronunciation objects from the speech information sequence based on the speech mask information of different pronunciation objects and the speech feature sequence.
[0273] Optionally, the processor can further execute program codes of the following steps: converting the speech features in the speech feature sequence according to a single-head attention mechanism to obtain the local speech information; converting the speech features in the speech feature sequence according to a linear attention mechanism to obtain the global speech information; and performing gating processing on the local speech information and the global speech information to obtain the gating processing result.
[0274] Optionally, the processor can further execute program codes of the following steps: performing convolution processing on the speech features in the speech feature sequence to obtain a speech feature matrix of a target dimension; converting the speech features in the speech feature sequence according to a linear attention mechanism to obtain the global speech information, comprising: converting the speech feature matrix according to the linear attention mechanism to obtain the global speech information.
[0275] Optionally, the processor can further execute program codes of the following steps: converting a block speech feature matrix of the speech feature matrix according to a single-head attention mechanism to obtain the local speech information.
[0276] Optionally, the processor can further execute program codes of the following steps: obtaining merged speech information between the global speech information and the local speech information; and performing gating processing on the merged speech information, the speech feature matrix and the speech feature sequence to obtain the gating processing result.
[0277] Optionally, the processor can further execute program codes of the following steps: performing convolution processing on the speech feature sequence for multiple times to obtain speech feature matrices of different target dimensions.
[0278] Optionally, the processor can further execute program codes of the following steps: performing normalization processing on the speech feature sequence to obtain a normalized speech result; performing encoding on the normalized speech result to obtain a speech encoding result; performing convolution processing on the speech encoding result, and converting the obtained convolution result to obtain a speech feature matrix of an original dimension; wherein the speech feature in the speech feature sequence is subjected to convolution processing to obtain a speech feature matrix of a target dimension, including: performing convolution processing on the speech feature matrix of the original dimension to obtain the speech feature matrix of the target dimension.
[0279] Optionally, the processor can further execute program codes of the following steps: performing convolution processing on the speech information sequence to obtain speech features of different pronunciation objects; performing linear processing on the speech features of different pronunciation objects to obtain the speech feature sequence.
[0280] Optionally, the processor can further execute program codes of the following steps: performing linear processing on the gating processing result, and performing convolution processing on the obtained linear processing result to obtain speech mask information of different pronunciation objects.
[0281] Optionally, the processor can further execute program codes of the following steps: obtaining a product result between the speech mask information of different pronunciation objects and the speech feature sequence; and determining the speech information output by different pronunciation objects based on the product result.
[0282] As an optional example, the processor can call information and application programs stored in the memory through the transmission device to execute the following steps: obtaining a speech information sequence, wherein the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different pronunciation objects; calling a speech separation model, wherein the speech separation model is obtained based on local attention mechanism and global attention mechanism; using the speech separation model to extract speech features of different pronunciation objects from the speech information sequence to obtain a speech feature sequence, and performing gating processing on the speech features in the speech feature sequence according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different pronunciation objects, and the information granularity of the local speech information is smaller than that of the global speech information; based on the gating processing result, obtaining speech mask information of different pronunciation objects, wherein the speech mask information is used to represent pronunciation attributes of the pronunciation objects; based on the speech mask information of different pronunciation objects and the speech feature sequence, separating speech information output by different pronunciation objects from the speech information sequence.
[0283] As an optional example, the processor can call information and application programs stored in the memory through the transmission device to perform the following steps: extracting speech features of different sound objects from the obtained speech information sequence to obtain a speech feature sequence, wherein the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different sound objects; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different sound objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound objects based on the gating processing result, wherein the speech mask information is used to represent the sound properties of the sound objects; separating the speech information output by different sound objects from the speech information sequence based on the speech mask information of different sound objects and the speech feature sequence; and playing the speech information output by different sound objects respectively.
[0284] As an optional example, the processor can call information and application programs stored in the memory through the transmission device to perform the following steps: extracting speech features of different sound objects from the obtained speech information sequence to obtain a speech feature sequence, wherein the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different sound objects; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different sound objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound objects based on the gating processing result, wherein the speech mask information is used to represent the sound properties of the sound objects; separating the speech information output by different sound objects from the speech information sequence based on the speech mask information of different sound objects and the speech feature sequence; and playing the speech information output by different sound objects respectively.
[0285] As an optional example, the processor can call the information and the application program stored in the memory through the transmission device to perform the following steps: obtaining a voice information sequence through calling a first interface, wherein the first interface comprises a first parameter, the parameter value of the first parameter is the voice information sequence, the voice information sequence comprises at least one voice information to be separated, and different voice information comes from different pronunciation objects; extracting the voice features of different pronunciation objects from the voice information sequence to obtain a voice feature sequence; performing gating processing on the voice features in the voice feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local voice information and global voice information of different pronunciation objects, and the information granularity of the local voice information is smaller than that of the global voice information; obtaining voice mask information of different pronunciation objects based on the gating processing result, wherein the voice mask information is used to represent the pronunciation attribute of the pronunciation object; separating the voice information output by different pronunciation objects from the voice information sequence based on the voice mask information of different pronunciation objects and the voice feature sequence; and outputting the voice information output by different pronunciation objects through calling a second interface, wherein the second interface comprises a second parameter, and the value of the second parameter is the voice information output by different pronunciation objects.
[0286] The embodiment of the present application performs gating processing on the voice features in the obtained voice information sequence according to a local attention mechanism and a global attention mechanism, and can obtain local voice information and global voice information of different pronunciation objects. Based on the gating processing, the requirements for the local attention mechanism and the global attention mechanism are greatly reduced, so that not only global information can be directly processed, but also smaller local features can be processed, thereby realizing the technical effect that the voice can be separated, and further solving the technical problem that the voice cannot be separated.
[0287] Those skilled in the art can understand that, Figure 18 The structure shown is only schematic, and the computer terminal A can also be a smart phone (such as a tablet computer, a palm computer, a mobile Internet device (MID), a PAD, and the like). Figure 18 The structure of the computer terminal A is not limited. For example, the computer terminal A can further comprise more or less components (such as a network interface, a display device, and the like), or have a different configuration from that shown. Figure 18 The structure of the computer terminal A is not limited. For example, the computer terminal A can further comprise more or less components (such as a network interface, a display device, and the like), or have a different configuration from that shown. The structure of the computer terminal A is not limited. For example, the computer terminal A can further comprise more or less components (such as a network interface, a display device, and the like), or have a different configuration from that shown.
[0288] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by instructing the terminal device related hardware through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0289] Embodiment 5
[0290] The embodiments of the present application also provide a computer readable storage medium. Optionally, in the embodiment, the computer readable storage medium can be used to save the program code executed by the speech separation method provided in the embodiment 1.
[0291] Optionally, in the embodiment, the computer readable storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0292] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: obtaining a speech information sequence, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different sound producing objects; extracting speech features of different sound producing objects from the speech information sequence to obtain a speech feature sequence; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different sound producing objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound producing objects based on the gating processing result, wherein the speech mask information is used to represent the sound producing properties of the sound producing objects; and separating the speech information output by different sound producing objects from the speech information sequence based on the speech mask information of different sound producing objects and the speech feature sequence
[0293] Optionally, the computer readable storage medium can also execute program code for performing the following steps: converting the speech features in the speech feature sequence according to a single-head attention mechanism to obtain the local speech information; converting the speech features in the speech feature sequence according to a linear attention mechanism to obtain the global speech information; and performing gating processing on the local speech information and the global speech information to obtain the gating processing result.
[0294] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing convolution processing on the speech features in the speech feature sequence to obtain a speech feature matrix of a target dimension; and converting the speech features in the speech feature sequence according to a linear attention mechanism to obtain global speech information, including: converting the speech feature matrix according to the linear attention mechanism to obtain the global speech information.
[0295] Optionally, the computer readable storage medium can further execute program codes of the following steps: converting the block speech feature matrix of the speech feature matrix according to a single-head attention mechanism to obtain local speech information.
[0296] Optionally, the computer readable storage medium can further execute program codes of the following steps: obtaining merged speech information between the global speech information and the local speech information; and performing gating processing on the merged speech information, the speech feature matrix, and the speech feature sequence to obtain a gating processing result.
[0297] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing multiple convolution processing on the speech feature sequence to obtain speech feature matrices of different target dimensions.
[0298] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing normalization processing on the speech feature sequence to obtain a normalized speech result; encoding the normalized speech result to obtain a speech encoding result; performing convolution processing on the speech encoding result, and converting the obtained convolution result to obtain a speech feature matrix of an original dimension; wherein the convolution processing on the speech features in the speech feature sequence to obtain the speech feature matrix of the target dimension includes: performing convolution processing on the speech feature matrix of the original dimension to obtain the speech feature matrix of the target dimension.
[0299] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing convolution processing on the speech information sequence to obtain speech features of different pronunciation objects; and performing linear processing on the speech features of the different pronunciation objects to obtain the speech feature sequence.
[0300] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing linear processing on the gating processing result, and performing convolution processing on the obtained linear processing result to obtain speech mask information of different pronunciation objects.
[0301] Optionally, the computer readable storage medium can further execute program codes of the following steps: obtaining a product result between the speech mask information of the different pronunciation objects and the speech feature sequence; and determining speech information output by the different pronunciation objects based on the product result.
[0302] As an optional example, the computer readable storage medium is configured to store program code for performing the following steps: obtaining a speech information sequence, wherein the speech information sequence comprises at least one speech information to be subjected to speech separation, and different speech information comes from different sound producing objects; calling a speech separation model, wherein the speech separation model is obtained based on local attention mechanism and global attention mechanism; using the speech separation model to extract speech features of different sound producing objects from the speech information sequence to obtain a speech feature sequence, and performing gating processing on the speech features in the speech feature sequence according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different sound producing objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound producing objects based on the gating processing result, wherein the speech mask information is used to represent the sound producing properties of the sound producing objects; and separating speech information output by different sound producing objects from the speech information sequence based on the speech mask information of different sound producing objects and the speech feature sequence.
[0303] As an optional example, the computer readable storage medium is configured to store program code for performing the following steps: from the obtained speech information sequence, extracting speech features of different sound producing objects to obtain a speech feature sequence, wherein the speech information sequence comprises at least one speech information to be subjected to speech separation, and different speech information comes from different sound producing objects; performing gating processing on the speech features in the speech feature sequence according to the local attention mechanism and the global attention mechanism to obtain a gating processing result, wherein the gating processing result comprises local speech information and global speech information of different sound producing objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound producing objects based on the gating processing result, wherein the speech mask information is used to represent the sound producing properties of the sound producing objects; and separating speech information output by different sound producing objects from the speech information sequence based on the speech mask information of different sound producing objects and the speech feature sequence; and playing the speech information output by different sound producing objects respectively.
[0304] As an optional example, the computer readable storage medium is configured to store program code for performing the following steps: extracting speech features of different sound objects from the obtained speech information sequence to obtain a speech feature sequence, wherein the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different sound objects; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different sound objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound objects based on the gating processing result, wherein the speech mask information is used to represent the sound properties of the sound objects; separating the speech information output by different sound objects from the speech information sequence based on the speech mask information of different sound objects and the speech feature sequence; and inputting the speech information output by different sound objects to a speech recognition end, wherein the speech information is used to be recognized by the speech recognition end.
[0305] As an optional example, the computer readable storage medium is configured to store program code for performing the following steps: obtaining a speech information sequence by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter is the speech information sequence, the speech information sequence includes at least one speech information to be subjected to speech separation, and different speech information comes from different sound objects; extracting speech features of different sound objects from the speech information sequence to obtain a speech feature sequence; performing gating processing on the speech features in the speech feature sequence according to a local attention mechanism and a global attention mechanism to obtain a gating processing result, wherein the gating processing result includes local speech information and global speech information of different sound objects, and the information granularity of the local speech information is smaller than that of the global speech information; obtaining speech mask information of different sound objects based on the gating processing result, wherein the speech mask information is used to represent the sound properties of the sound objects; separating the speech information output by different sound objects from the speech information sequence based on the speech mask information of different sound objects and the speech feature sequence; and outputting the speech information output by different sound objects by calling a second interface, wherein the second interface includes a second parameter, and the value of the second parameter is the speech information output by different sound objects.
[0306] The above-mentioned serial numbers of the embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0307] In the above-mentioned embodiments of the application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0308] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented in other manners. For example, the described embodiments of the apparatus are merely schematic, and the division of units is merely logical function division, and there can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, and electrical or other forms.
[0309] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0310] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.
[0311] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially or substantially, or all or part of the technical solutions that make contributions to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various other media that can store program codes.
[0312] The above descriptions are merely preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be regarded as the protection scope of the present application.
Claims
1. A speech separation method, characterized in that, include: Acquire a speech information sequence, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different speech objects; The speech features of different pronunciation objects are extracted from the speech information sequence to obtain a speech feature sequence; The speech features in the speech feature sequence are gating processed according to local attention mechanism and global attention mechanism to obtain gating processing result, wherein the gating processing result includes local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; Based on the gating processing result, speech mask information of different pronunciation objects is obtained, wherein the speech mask information is used to represent the pronunciation attributes of the pronunciation object; Based on the speech mask information of different pronunciation objects and the speech feature sequence, the speech information output by different pronunciation objects is separated from the speech information sequence.
2. The method according to claim 1, characterized in that, The local attention mechanism includes a single-head attention mechanism, and the global attention mechanism includes a linear attention mechanism. The speech features in the speech feature sequence are gating processed according to both the local and global attention mechanisms to obtain the gating result, including: The speech features in the speech feature sequence are transformed according to the single-head attention mechanism to obtain the local speech information; The speech features in the speech feature sequence are transformed according to the linear attention mechanism to obtain the global speech information; The local speech information and the global speech information are subjected to gating processing to obtain the gating processing result.
3. The method according to claim 2, characterized in that, The method further includes: The speech features in the speech feature sequence are convolved to obtain a speech feature matrix of the target dimension; Transforming the speech features in the speech feature sequence according to the linear attention mechanism to obtain the global speech information includes: transforming the speech feature matrix according to the linear attention mechanism to obtain the global speech information.
4. The method according to claim 3, characterized in that, The speech features in the speech feature sequence are transformed according to the single-head attention mechanism to obtain the local speech information, including: The segmented speech feature matrix of the speech feature matrix is transformed according to the single-head attention mechanism to obtain the local speech information.
5. The method according to claim 4, characterized in that, Gating processing is performed on the local speech information and the global speech information to obtain the gating processing result, including: Obtain the merged speech information between the global speech information and the local speech information; The merged speech information, the speech feature matrix, and the speech feature sequence are subjected to gating processing to obtain the gating processing result.
6. The method according to claim 3, characterized in that, The speech feature sequence is convolved to obtain a speech feature matrix of the target dimension, including: The speech feature sequence is subjected to convolution processing multiple times to obtain speech feature matrices of different target dimensions.
7. The method according to claim 3, characterized in that, The method further includes: The speech feature sequence is normalized to obtain the normalized speech result; The normalized speech result is encoded to obtain the speech coding result; The speech coding result is convolved, and the resulting convolution result is transformed to obtain the original dimension speech feature matrix; Specifically, performing convolution processing on the speech features in the speech feature sequence to obtain a speech feature matrix of the target dimension includes: performing convolution processing on the original dimension speech feature matrix to obtain the speech feature matrix of the target dimension.
8. The method according to claim 1, characterized in that, Speech features of different pronunciation objects are extracted from the speech information sequence to obtain a speech feature sequence, including: The speech information sequence is convolved to obtain the speech features of different pronunciation objects; The speech feature sequence is obtained by linearly processing the speech features of different pronunciation objects.
9. The method according to any one of claims 1 to 8, characterized in that, Based on the gating processing results, speech mask information for different speech objects is obtained, including: The gating processing result is linearly processed, and the resulting linear processing result is convolved to obtain speech mask information for different speech objects.
10. The method according to any one of claims 1 to 8, characterized in that, Based on the speech mask information of different pronunciation objects and the speech feature sequence, the speech information output by different pronunciation objects is separated from the speech information sequence, including: Obtain the product of the speech mask information of different pronunciation objects and the speech feature sequence; The product result determines the speech information output by different pronunciation objects.
11. A speech separation method, characterized in that, include: Acquire a speech information sequence, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different speech objects; The speech separation model is invoked, wherein the speech separation model is obtained by training based on local attention mechanism and global attention mechanism; Using the speech separation model, speech features of different pronunciation objects are extracted from the speech information sequence to obtain a speech feature sequence. The speech features in the speech feature sequence are then gating processed according to local attention mechanism and global attention mechanism to obtain gating processing result. The gating processing result includes local speech information and global speech information of different pronunciation objects. The information granularity of the local speech information is smaller than that of the global speech information. Based on the gating processing result, speech mask information of different pronunciation objects is obtained, wherein the speech mask information is used to represent the pronunciation attributes of the pronunciation object; Based on the speech mask information of different pronunciation objects and the speech feature sequence, the speech information output by different pronunciation objects is separated from the speech information sequence.
12. A speech separation method, characterized in that, include: From the acquired speech information sequence, speech features of different pronunciation objects are extracted to obtain a speech feature sequence, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different pronunciation objects; The speech features in the speech feature sequence are gating processed according to local attention mechanism and global attention mechanism to obtain gating processing result, wherein the gating processing result includes local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; Based on the gating processing result, speech mask information of different pronunciation objects is obtained, wherein the speech mask information is used to represent the pronunciation attributes of the pronunciation object; Based on the speech mask information of different pronunciation objects and the speech feature sequence, the speech information output by different pronunciation objects is separated from the speech information sequence; The speech information output by different speech objects is played respectively.
13. A speech separation method, characterized in that, include: From the acquired speech information sequence, speech features of different pronunciation objects are extracted to obtain a speech feature sequence, wherein the speech information sequence includes at least one speech information to be separated, and different speech information comes from different pronunciation objects; The speech features in the speech feature sequence are gating processed according to local attention mechanism and global attention mechanism to obtain gating processing result, wherein the gating processing result includes local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; Based on the gating processing result, speech mask information of different pronunciation objects is obtained, wherein the speech mask information is used to represent the pronunciation attributes of the pronunciation object; Based on the speech mask information of different speech objects and the speech feature sequence, the speech information output by different speech objects is separated from the speech information sequence; The speech information output by different speech objects is input to the speech recognition terminal, wherein the speech information is used for recognition by the speech recognition terminal.
14. A speech separation method, characterized in that, include: A speech information sequence is obtained by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter is the speech information sequence, and the speech information sequence includes at least one speech information to be separated, and different speech information comes from different pronunciation objects; The speech features of different pronunciation objects are extracted from the speech information sequence to obtain a speech feature sequence; The speech features in the speech feature sequence are gating processed according to local attention mechanism and global attention mechanism to obtain gating processing result, wherein the gating processing result includes local speech information and global speech information of different speech objects, and the information granularity of the local speech information is smaller than that of the global speech information; Based on the gating processing result, speech mask information of different pronunciation objects is obtained, wherein the speech mask information is used to represent the pronunciation attributes of the pronunciation object; Based on the speech mask information of different speech objects and the speech feature sequence, the speech information output by different speech objects is separated from the speech information sequence; The second interface outputs the speech information of different speech objects by calling the second interface, wherein the second interface includes a second parameter, and the value of the second parameter is the speech information output by different speech objects.
Citation Information
Patent Citations
Face micro-expression recognition method in video image sequence
CN113496217A
Method, medium, and apparatus for extracting target sound from mixed sound
US20090097670A1