Voice signal acquisition method and device, computer device, and storage medium

By performing various temporal resolution feature extractions and attention mechanisms on the speech signal, the distortion problem of the target speech signal in multi-speaker speech environments is solved, and more accurate speech signal acquisition is achieved.

CN116564288BActive Publication Date: 2025-12-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210095158.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-12-19
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

Existing speaker extraction techniques suffer from distortion in target speech signals acquired in multi-speaker speech environments, with residual buzzing or excessive suppression of the speaker's voice, resulting in inaccurate signals.

Method used

By extracting features from speech signals at various time resolutions, processing multiple feature channels using an attention mechanism, and combining these features with the speech characteristics of the target object, a more accurate target speech signal can be obtained.

Benefits of technology

It effectively suppresses irrelevant feature channels, improves the accuracy of speech signals, and obtains more accurate speech signals of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116564288B_ABST
    Figure CN116564288B_ABST
Patent Text Reader

Abstract

The application provides a speech signal acquisition method and device, computer equipment and a storage medium, belonging to the multimedia technical field. The method comprises: extracting features of a speech signal according to multiple time resolutions to obtain multiple first speech features; for any first speech feature, processing multiple feature channels in the first speech feature based on an attention mechanism to obtain a second speech feature corresponding to the first speech feature; and acquiring a speech signal of a target object from the speech signal based on the multiple first speech features, the multiple second speech features and a speech feature of the target object. The above scheme can suppress the feature channels of irrelevant features in the speech feature by processing the multiple feature channels in the speech feature based on the attention mechanism, improve the performance of the speech feature, and further acquire a more accurate speech signal of the target object from the speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multimedia, and particularly relates to a speech signal acquisition method and device, a computer device and a storage medium. BACKGROUND

[0002] Speaker extraction technology is a technology for acquiring target speech signals, and can selectively extract speech signals of a speaker of interest from a multi-speaker speech environment. For example, in a multi-person online conference scenario, target speech signals of a main speaker can be acquired from played speech signals through speaker extraction technology.

[0003] However, the target speech signals extracted through speaker extraction technology may have distortion problems, that is, there may be residual squeaking or excessive suppression of the voice of the main speaker in the acquired target speech signals, so that the acquired target speech signals are not accurate enough. SUMMARY

[0004] Embodiments of the present application provide a speech signal acquisition method and device, a computer device and a storage medium, which can acquire more accurate speech signals of a target object from speech signals. The technical solutions are as follows:

[0005] In one aspect, a speech signal acquisition method is provided according to an embodiment of the present application, and the method comprises:

[0006] characteristic extraction of the speech signal according to a plurality of time resolutions to obtain a plurality of first speech characteristics, the first speech characteristics corresponding to the time resolutions, and the first speech characteristics comprising a plurality of feature channels;

[0007] for any first speech characteristic, processing a plurality of feature channels in the first speech characteristic based on an attention mechanism to obtain a second speech characteristic corresponding to the first speech characteristic;

[0008] acquiring speech signals of a target object from the speech signals based on the plurality of first speech characteristics, the plurality of second speech characteristics and speech characteristics of the target object.

[0009] In another aspect, a speech signal acquisition device is provided according to an embodiment of the present application, and the device comprises:

[0010] a first extraction module configured to perform characteristic extraction of a speech signal according to a plurality of time resolutions to obtain a plurality of first speech characteristics, the first speech characteristics corresponding to the time resolutions, and the first speech characteristics comprising a plurality of feature channels;

[0011] The first processing module is configured to, for any first speech feature, process multiple feature channels in the first speech feature based on an attention mechanism to obtain a second speech feature corresponding to the first speech feature.

[0012] The first obtaining module is configured to obtain a speech signal of the target object from the speech signal based on the multiple first speech features, the multiple second speech features, and a speech feature of the target object.

[0013] In some embodiments, the first extraction module comprises:

[0014] The first extraction unit is configured to, for any speech channel, perform feature extraction on a signal of the speech channel in multiple time resolutions to obtain multiple speech features of the speech channel.

[0015] The second extraction unit is configured to perform spatial feature extraction on the speech features of the multiple speech channels to obtain a spatial information feature, the spatial information feature being used to represent a spatial relationship between the speech features of the multiple speech channels.

[0016] The summing unit is configured to add the multiple speech features of a target speech channel and the spatial information feature respectively to obtain the multiple first speech features, the target speech channel being a channel corresponding to the target object.

[0017] In some embodiments, the first obtaining module comprises:

[0018] The determining unit is configured to determine a mask feature based on the multiple second speech features and the speech feature of the target object.

[0019] The third extraction unit is configured to extract a target speech feature from the multiple first speech features based on the mask feature.

[0020] The decoding unit is configured to decode the target speech feature to obtain the speech signal of the target object.

[0021] In some embodiments, the determining unit is configured to perform normalization on the multiple second speech features respectively, and determine the mask feature based on the normalized multiple second speech features and the speech feature of the target object.

[0022] In some embodiments, the third extraction unit is configured to splice the multiple first speech features to obtain spliced features, and perform element multiplication on the mask feature and the spliced features to obtain the target speech feature.

[0023] In some embodiments, the apparatus further comprises:

[0024] The second acquisition module is configured to acquire a reference voice signal, the reference voice signal being a voice signal of the target object.

[0025] The second extraction module is configured to perform feature extraction on the reference voice signal according to the multiple time resolutions to obtain multiple third voice features, the third voice features corresponding to the time resolutions.

[0026] The second processing module is configured to perform feature embedding processing on the multiple third voice features to obtain a voice feature of the target object.

[0027] In some embodiments, the apparatus further includes:

[0028] The third acquisition module is configured to, for any sub-mask feature, acquire at least two sub-mask features that are time-adjacent to the sub-mask feature.

[0029] The fusion module is configured to fuse the at least two sub-mask features and the sub-mask feature to obtain a fused sub-mask feature.

[0030] In another aspect, a computer device is provided, which includes a processor and a memory, the memory being configured to store at least one piece of computer program, the at least one piece of computer program being loaded and executed by the processor to implement the voice signal acquisition method in the embodiments of the present application.

[0031] In another aspect, a computer readable storage medium is provided, which stores at least one piece of computer program, the at least one piece of computer program being loaded and executed by a processor to implement the voice signal acquisition method in the embodiments of the present application.

[0032] In another aspect, a computer program product is provided, which includes computer program code stored in a computer readable storage medium, a processor of a computer device reading the computer program code from the computer readable storage medium, and the processor executing the computer program code to cause the computer device to perform the voice signal acquisition method provided in various optional implementations of each aspect.

[0033] The embodiments of the present application provide a voice signal acquisition scheme, which can suppress irrelevant feature channels in a voice feature, improve the performance of the voice feature, and further acquire a more accurate voice signal of a target object from a voice signal by processing multiple feature channels in the voice feature based on an attention mechanism. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.

[0035] Figure 1 is a schematic diagram of an implementation environment of a voice signal acquisition method according to an embodiment of the present application;

[0036] Figure 2 is a flowchart of a voice signal acquisition method according to an embodiment of the present application;

[0037] Figure 3 is a flowchart of another voice signal acquisition method according to an embodiment of the present application;

[0038] Figure 4 is a schematic diagram of a single-voice-channel voice signal acquisition method according to an embodiment of the present application;

[0039] Figure 5 is a schematic diagram of a multi-voice-channel voice signal acquisition method according to an embodiment of the present application;

[0040] Figure 6 is a schematic diagram of a sub-mask feature fusion based on context information according to an embodiment of the present application;

[0041] Figure 7 is a schematic diagram of a speech spectrogram of a target object voice signal according to an embodiment of the present application;

[0042] Figure 8 is a structural schematic diagram of a voice signal acquisition device according to an embodiment of the present application;

[0043] Figure 9 is a structural schematic diagram of another voice signal acquisition device according to an embodiment of the present application;

[0044] Figure 10 is a structural block diagram of a terminal according to an embodiment of the present application;

[0045] Figure 11 is a structural schematic diagram of a server according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of the present application more clear, the embodiments of the present application will be further described in detail with reference to the drawings.

[0047] The terms "first", "second", and the like are used to distinguish between the same or similar items or elements having essentially the same function, and there is no logical or chronological dependency between the "first", "second", and "n", nor is the number and execution order limited.

[0048] The term "at least one" in the present application means one or more, and the meaning of "a plurality" is two or more.

[0049] In the detailed description of the present application, data related to voice signals and the like are involved, and when the above embodiments of the present application are applied to specific products or technologies, the user's permission or consent is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region.

[0050] For ease of understanding, the terms involved in the present application are explained as follows.

[0051] Cloud conference: a kind of efficient, convenient and low-cost conference form based on cloud computing technology. Users only need to use the Internet interface and simple operation to quickly and efficiently share voice, data files and video with teams and customers around the world, and the complex technology of data transmission and processing in the conference is operated by the cloud conference service provider.

[0052] At present, cloud conference mainly focuses on the service content of SaaS (Software as a Service) mode, including telephone, network, video and other service forms, and video conference based on cloud computing is called cloud conference. In the era of cloud conference, data transmission, processing and storage are all handled by the computer resources of the video conference manufacturer, and users no longer need to purchase expensive hardware and install complicated software. They only need to open a browser and log in to the corresponding interface to conduct efficient remote conference.

[0053] The cloud conference system supports multi-server dynamic cluster deployment and provides multiple high-performance servers, greatly improving the stability, security and availability of the conference. In recent years, video conference has been widely used in transportation, finance, operators, education, enterprises and other fields because it can greatly improve communication efficiency, continuously reduce communication costs and upgrade internal management. There is no doubt that video conference using cloud computing has stronger attraction in convenience, speed and ease of use, and will stimulate a new high tide of video conference application.

[0054] Speaker extraction technology: a technology for obtaining target voice signals, which can selectively extract the voice signals of the speaker of interest from a multi-speaker voice environment.

[0055] SpEx+ model: a network model based on the time-domain features of the speech signal to extract the speech signal of the speaker.

[0056] Attention mechanism: an information processing method that selectively focuses on part of all information while ignoring other visible information.

[0057] CNN (Convolutional Neural Network): a class of feedforward neural networks containing convolutional calculations and having a deep structure, which is one of the representative algorithms of deep learning. CNN includes one-dimensional convolutional neural networks and two-dimensional convolutional neural networks. The two-dimensional convolutional network performs sliding window operations on a feature map in the width and height directions and multiplies and sums the corresponding positions. The one-dimensional convolutional neural network only performs sliding window operations in the width or height direction and multiplies and sums.

[0058] Contextual mechanism: a method of processing target information based on the context information of the target information, which is the information in the neighborhood of the target information.

[0059] ResNet (Residual Network): a new CNN architecture with skip connections and batch normalization, which can train a 152-layer neural network. The residual function is learned by stacking layer sets, which is easier to optimize and can greatly deepen the network.

[0060] Feature embedding: a method of converting data into fixed-size feature representations for processing and calculation. For example, for extracting the speech signal of the speaker from the mixed speech signal, the speech signal can be converted into a feature vector, so that the reference speech signal from the same speaker has a small distance from the original vector.

[0061] IPD (Interaural Phase Difference): the difference in phase of sound waves at both ears caused by the time difference of a periodic sound wave reaching both ears.

[0062] CD (Channel Decorrelation): a method of constructing a latent embedding space to learn dimension-level discriminative features between multi-channel input mixed encodings through the spatial structure of a multi-channel microphone array.

[0063] Beamforming: a signal processing technique used in sensor arrays for directional signal transmission or reception. This technique can enhance the signal at a certain angle (target user) and weaken the signal at another angle (non-target user or obstacle), while achieving spatial selectivity at the sending and receiving ends.

[0064] TCN (Temporal Convolutional Networks): a convolutional neural network that can be used for time series data processing.

[0065] ReLU (Linear rectification function): an activation function commonly used in artificial neural networks, usually referring to a nonlinear function represented by a ramp function and its variants.

[0066] The speech signal acquisition method provided by the embodiments of the present application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server. In the following, taking the computer device as a server as an example, the implementation environment of the speech signal acquisition method provided by the embodiments of the present application is introduced, Figure 1 is a schematic diagram of the implementation environment of the speech signal acquisition method provided by the embodiments of the present application. Referring to Figure 1 , the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.

[0067] In some embodiments, the terminal 101 is a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, and a vehicle-mounted terminal, but is not limited thereto. The terminal 101 is installed and runs an application program supporting the acquisition of speech signals. The application program is any one of a variety of application programs such as a multimedia application program, a social application program, or a conference application program. The terminal 101 is a terminal used by a user, and the user uses the terminal 101 to acquire speech signals in the surrounding environment. It should be noted that the number of terminals can be more or less. For example, the number of terminals is one, or the number of terminals is dozens or hundreds, or more. The number of terminals and the type of devices are not limited in the embodiments of the present application.

[0068] In some embodiments, the server 102 is a stand-alone physical server, can also be a server cluster or distributed system composed of multiple physical servers, and can also be a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN (Content Delivery Network), and big data and artificial intelligence platform. The server 102 is used to provide background services for an application program supporting voice signal acquisition. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing.

[0069] In some embodiments, in a cloud conference scenario, the server 102 can acquire a target voice signal from a voice signal by using the voice signal acquisition method provided in the embodiments of the present application. The terminal 101 acquires a voice signal in the surrounding environment, and then sends the voice signal to the server 102. After receiving the voice signal, the server 102 acquires a target voice signal from the voice signal by using the voice signal acquisition method, and then sends the target voice signal to other terminals 101, which play the target voice signal.

[0070] For example, in a multi-person online conference scenario, the terminal 101 acquires a voice signal, which includes a voice signal of a main speaker, voice signals of other conference participants, and a noise signal. The terminal 101 sends the voice signal to the server 102. After receiving the voice signal, the server 102 acquires the voice signal of the main speaker from the voice signal by using the voice signal acquisition method provided in the embodiments of the present application, and then sends the voice signal of the main speaker to the terminals 101 of the other conference participants, which play the voice of the main speaker to the other conference participants.

[0071] In some embodiments, in an offline conference scenario, the terminal 101 can also acquire a target voice signal from a voice signal locally by using the voice signal acquisition method provided in the embodiments of the present application, and then deliver the target voice signal to a user.

[0072] For example, in a live conference scenario, the terminal 101 acquires a voice signal, which includes a voice signal of a main speaker and a noise signal. The terminal 101 can acquire the voice signal of the main speaker from the voice signal by using the voice signal acquisition method provided in the embodiments of the present application, and then play the voice of the main speaker to the conference participants in the live conference by broadcasting.

[0073] Figure 2 is a flowchart of a speech signal acquisition method according to an embodiment of the present application, referring to Figure 2 In the embodiments of the present application, a server is taken as an example for illustration. The speech signal acquisition method comprises the following steps:

[0074] 201. The server extracts features of the speech signal according to multiple time resolutions, to obtain multiple first speech features, each first speech feature corresponding to a time resolution, and each first speech feature comprising multiple feature channels.

[0075] In the embodiments of the present application, the speech signal is a mixed speech signal, comprising a speech signal of a target object and a background sound signal. The background sound signal can be a speech signal of another object, or a noise signal, or a mixed signal of a speech signal of another object and a noise signal, which is not limited in the embodiments of the present application. The time resolution refers to the length of a time window during feature extraction, and the time window can divide the speech signal into multiple speech signal segments. The lower the time resolution of the speech signal, the longer the duration of each speech signal segment, and vice versa, the higher the time resolution of the speech signal, the shorter the duration of each speech signal segment. The server extracts features of the speech signal based on multiple time resolutions, to obtain multiple first speech features. Each first speech feature corresponds to a time resolution, and different first speech features correspond to different time resolutions. The number of elements of the first speech features extracted based on different time resolutions is different. The higher the time resolution, the more the number of elements of the first speech features, and the lower the time resolution, the less the number of elements of the first speech features. The server can extract features of the speech signal based on a one-dimensional convolutional neural network, and each first speech feature extracted comprises multiple feature channels. The number of feature channels included in each speech feature is the same, and the number of feature channels is equal to the number of convolution kernels of the one-dimensional convolutional neural network. The size of the convolution kernel of different one-dimensional convolutional neural networks is different, and the size of multiple different convolution kernels of the same one-dimensional convolutional neural network is the same.

[0076] 202. For any first speech feature, the server processes multiple feature channels in the first speech feature based on an attention mechanism, to obtain a second speech feature corresponding to the first speech feature.

[0077] In the embodiment of the present application, the attention mechanism is an information processing method that selectively focuses on part of the information in multiple information while ignoring other visible information. Since the voice signal includes the voice signal of the target object and the background sound signal, the first voice feature includes the feature of the voice signal of the target object and the feature of the background sound signal. For any first voice feature, the server can weight the multiple feature channels of the first voice feature in the dimension of the feature channel based on the attention mechanism, suppress the features of irrelevant feature channels, and thus improve the performance of different feature channels, wherein the attention mechanism is a channel attention mechanism.

[0078] 203. The server obtains the voice signal of the target object from the voice signal based on the multiple first voice features, the multiple second voice features, and the voice feature of the target object.

[0079] In the embodiment of the present application, the voice feature of the target object can be a voiceprint feature, a frequency feature, or an amplitude feature of the target object, and the embodiment of the present application does not limit it. The server processes the multiple first voice features respectively in the manner of step 202, and can obtain multiple second voice features, which correspond one-to-one to the first voice features. The second voice features also include the features of the voice signal of the target object and the background sound signal. The server takes the voice feature of the target object as a reference, and can obtain the feature of the voice signal of the target object from the first voice feature, so as to obtain the voice signal of the target object from the voice signal.

[0080] The scheme provided in the embodiment of the present application can suppress the irrelevant feature channels in the voice feature, improve the performance of the voice feature, and thus obtain more accurate voice signal of the target object from the voice signal.

[0081] The above Figure 2 The above Figure 3 is a flowchart of another voice signal acquisition method provided by the embodiment of the present application, see Figure 3 In the embodiment of the present application, a server is taken as an example for illustration. The voice signal acquisition method includes the following steps:

[0082] 301. The server performs feature extraction on the voice signal according to multiple time resolutions, and obtains multiple first voice features, wherein the first voice feature corresponds to a time resolution, and the first voice feature includes multiple feature channels.

[0083] In the embodiments of the present application, the voice signal is a single voice channel voice signal or a mixed signal of multiple voice channels. The single voice channel means that the voice signal comes from the same audio acquisition device, and the multiple voice channels means that the voice signal comes from different audio acquisition devices. The audio acquisition device can be a microphone, a walkie-talkie or a recorder, etc., which is not limited in the embodiments of the present application.

[0084] In some embodiments, the server can employ a convolutional neural network to extract the first voice features. The feature extraction process is that the server employs multiple one-dimensional convolutional neural networks to respectively convolve the voice signal, to obtain multiple first voice features. The multiple one-dimensional convolutional neural networks have different sizes of convolution kernels. Since the sizes of the convolution kernels are different, the lengths of the time windows used by the one-dimensional convolutional neural networks in the feature extraction of the voice signal are different, see step 201, which is not described here. Based on the different sizes of the convolution kernels, convolution operations on the voice signal under different time resolutions can be realized. It can be understood that each first voice feature obtained based on the above process corresponds to a time resolution. In addition, each first voice feature includes multiple feature channels, and the number of the multiple feature channels is the same as the number of the convolution kernels in the one-dimensional convolutional neural network, which can be 128, 256 or 512, which is not limited in the embodiments of the present application.

[0085] In some embodiments, the voice signal is a single voice channel signal, and the voice signal comes from a single audio acquisition device. The process of the server obtaining multiple first voice features is that the server obtains the voice signal collected based on the single audio acquisition device, and then performs feature extraction on the voice signal according to multiple time resolutions to obtain multiple first voice features.

[0086] For example, Figure 4 is a schematic diagram of a method for obtaining a single voice channel voice signal according to the embodiments of the present application. Referring to Figure 4 , the voice signal is collected based on a microphone worn by a target object, a handheld walkie-talkie or other audio acquisition device, and the audio acquisition device is bound with an object identifier of the target object. The server inputs the voice signal into a speech encoder, and the speech encoder is used to extract voice features of the voice signal, and the speech encoder includes three one-dimensional convolutional neural networks with different sizes of convolution kernels. The server extracts features of the input voice signal by the speech encoder using one-dimensional convolutional neural networks with different sizes of convolution kernels to obtain multiple first voice features.

[0087] In some embodiments, the voice signal comprises a multi-voice channel signal, i.e., the voice signal is mixed from signals from different audio acquisition devices. Accordingly, the process of obtaining the first voice features by the server can be implemented through the following three steps (1) to (3).

[0088] (1) For any voice channel, the server extracts features from the signal of the voice channel at multiple time resolutions to obtain multiple voice features of the voice channel.

[0089] The principle of this step is the same as that of obtaining the first voice features of the single voice channel signal described above, and will not be repeated here.

[0090] (2) The server extracts spatial features from the voice features of the multiple voice channels to obtain spatial information features.

[0091] The spatial information features are used to represent the spatial relationship between the voice features of the multiple voice channels. The server can obtain the spatial information features through IPD technology, CD technology, or Beamforming technology, etc., which are not limited by the embodiments of the present application.

[0092] In addition, the server can also obtain the spatial information features through the following formula one.

[0093] Formula one:

[0094] W p =p(W1,…,W n )

[0095] Wherein, W p represents the spatial information features, W1 represents the voice features of the first voice channel, W n represents the voice features of the nth voice channel, and the function p(·) represents the spatial feature extraction operation.

[0096] (3) The server adds the multiple voice features of the target voice channel to the spatial information features respectively to obtain the multiple first voice features.

[0097] Wherein, the target voice channel is the channel corresponding to the target object, which can be any voice channel in the multiple voice channels.

[0098] The server can obtain the first voice features through the following formula two.

[0099] Formula two:

[0100] Y=W p +W1

[0101] Wherein, Y represents the first voice features, W pW1 represents the spatial information features, and W1 represents the speech features of the target speech channel.

[0102] For example, Figure 5 This is a schematic diagram of a method for acquiring multi-channel voice signals according to an embodiment of this application. See also... Figure 5 The server inputs speech signals from multiple speech channels into multiple speech encoders. Each speech channel corresponds to one speech encoder, and each speech encoder includes three one-dimensional convolutional neural networks with different kernel sizes; each kernel size corresponds to a different temporal resolution. For any given speech channel, the server uses the speech encoder corresponding to that channel to extract features from the signal at multiple temporal resolutions, obtaining multiple speech features for that channel. The server can then concatenate these multiple speech features to obtain the speech features for that channel. Next, the server inputs the speech features from multiple speech channels into a preprocessing module, which performs spatial feature extraction on these features to obtain spatial information features. Figure 5 An example is shown where the speech signal 1 originates from the target speech channel. Then, the server adds multiple speech features and spatial information features of the target speech channel to obtain multiple first speech features.

[0103] 302. For any first speech feature, the server processes multiple feature channels in the first speech feature based on an attention mechanism to obtain the second speech feature corresponding to the first speech feature.

[0104] In this embodiment, the server introduces a channel attention mechanism to weight multiple feature channels along the feature channel dimension, increasing the features of important feature channels and suppressing the features of irrelevant feature channels, thereby providing the performance capabilities of different feature channels and obtaining the second speech feature.

[0105] For example, see continue. Figure 4 or Figure 5 The server processes multiple first speech features separately based on an attention mechanism to obtain multiple second speech features.

[0106] In order to more accurately extract the speech signal of the target object from the speech signal, it is necessary to accurately extract the features of the target object's speech signal from the features of the speech signal. Therefore, it is necessary to obtain the speech features of the target object as a reference in order to accurately extract the features of the target object's speech signal from the features of the speech signal.

[0107] 303. The server obtains the speech features of the target object.

[0108] In the embodiments of the present application, the voice feature of the target object can be a voiceprint feature, a frequency feature, or an amplitude feature, etc., and the embodiments of the present application do not limit this. The server can extract the voice feature of the target object from the plurality of first voice features by taking the voice feature of the target object as a reference.

[0109] In some embodiments, the server can pre-acquire the voice feature of the target object and store the voice feature of the target object in a database. The server can directly acquire the voice feature of the target object from the database. By pre-acquiring the voice feature of the target object, the voice feature of the target object can be directly acquired when used, and the efficiency of acquiring the voice feature of the target object is high, thereby improving the efficiency of acquiring the voice signal of the target object.

[0110] In some embodiments, the server can acquire the voice signal of the target object in real time, and then determine the voice feature of the target object based on the voice signal of the target object. The process of the server acquiring the voice feature of the target object can be implemented through the following three steps (1)-(3).

[0111] (1) The server acquires a reference voice signal.

[0112] The reference voice signal is the voice signal of the target object. The reference voice signal can be obtained from the historical voice of the target object, or can be obtained from the voice of the target object collected on site. In some embodiments, the reference voice signal is a pure voice signal of the target object, that is, only the voice signal of the target object is included, and the voice signal of other objects is not included.

[0113] (2) The server extracts features from the reference voice signal according to a plurality of time resolutions to obtain a plurality of third voice features.

[0114] The plurality of time resolutions are the same as the plurality of time resolutions in step 301. That is, the server can use the same principle as that of extracting the first voice feature in step 301 to extract features from the reference voice signal based on the plurality of time resolutions.

[0115] (3) The server performs feature embedding processing on the plurality of third voice features to obtain the voice feature of the target object.

[0116] For example, continuing to refer to Figure 4 Or Figure 5The server inputs the reference voice signal into a speaker encoder. The speaker encoder is used to extract the voice feature of the target object. The speaker encoder sequentially includes a Layer Norm, a full connection layer, four residual network layers, a full connection layer, and an average pooling layer. For any third voice feature, the server first inputs the third voice feature into the Layer Norm and the full connection layer, and normalizes the plurality of third voice features in the feature channel direction. Then, the server inputs the normalized third voice feature into the four residual network layers in sequence, and then inputs the full connection layer and the average pooling layer to complete the feature embedding processing, thereby obtaining the voice feature of the target object.

[0117] It should be noted that this step can be executed before steps 301 and 302, after steps 301 and 302, or simultaneously with steps 301 and 302, and the embodiments of the present application do not limit this.

[0118] 304. The server determines a mask feature based on the plurality of second voice features and the voice feature of the target object, the mask feature including a plurality of sub-mask features.

[0119] In the embodiments of the present application, the server normalizes the plurality of second voice features respectively. Then, the server can determine the mask feature based on the normalized plurality of second voice features and the voice feature of the target object.

[0120] For example, continuing to refer to Figure 4 Or Figure 5The server inputs the normalized plurality of second speech features and the speech feature of the target object into a speaker extractor. The speaker extractor comprises a regularization layer, a fully connected layer, four TCN layers, a one-dimensional CNN, and a ReLU activation unit connected in sequence. First, the server inputs the plurality of second speech features into the regularization layer and the fully connected layer to normalize the plurality of second speech features. The server inputs the normalized plurality of second speech features and the speech feature of the target object into the first TCN layer. Then, the server inputs the output feature of the first TCN layer and the speech feature of the target object into the second TCN layer. Then, the server inputs the output feature of the second TCN layer and the speech feature of the target object into the third TCN layer. Then, the server inputs the output feature of the third TCN layer and the speech feature of the target object into the fourth TCN layer. Then, the server inputs the output feature of the fourth TCN layer into the one-dimensional convolutional neural network with different convolution kernel sizes. Finally, the server outputs the mask feature through the ReLU activation unit. The mask feature comprises a plurality of sub-mask features, and the number of the sub-mask features is inversely proportional to the size of the convolution kernel, that is, the longer the length of the convolution kernel, the fewer the number of the sub-mask features, and the shorter the length of the convolution kernel, the more the number of the sub-mask features.

[0121] 305、For any sub-mask feature, the server processes the sub-mask feature.

[0122] In the embodiment of the present application, for any sub-mask feature, the server can obtain at least two sub-mask features that are time-adjacent to the sub-mask feature, which can be referred to as context information of the sub-mask feature. Then, the server fuses the at least two sub-mask features with the sub-mask feature to obtain a fused sub-mask feature. The server performs the above fusion operation on each sub-mask feature in the mask feature to obtain a plurality of fused sub-mask features.

[0123] The server can process the sub-mask feature in the mask feature through the following formula three.

[0124] Formula three:

[0125] M' = f(M, C)

[0126] Wherein, M' represents the mask feature after processing, M represents the mask feature before processing, and C represents the number of sub-mask features that are time-adjacent to the sub-mask feature. C can be 1, 2, or 3, and the present application does not limit this.

[0127] For example, when C is 1, the server obtains two sub-mask features that are time-adjacent to the sub-mask feature. Figure 6is a schematic diagram of fusing sub-mask features based on context information according to an embodiment of the present application. Referring to Figure 6 The mask features before purification include a plurality of sub-mask features. For a first sub-mask feature, the server acquires a second sub-mask feature and a third sub-mask feature adjacent in time sequence to the first sub-mask feature. The server fuses the first sub-mask feature, the second sub-mask feature, and the third sub-mask feature to obtain a fused first sub-mask feature. When C is 2, the server acquires four sub-mask features adjacent in time sequence to the sub-mask feature. The four sub-mask features are two sub-mask features before the sub-mask feature in time sequence and two sub-mask features after the sub-mask feature in time sequence.

[0128] The server can more accurately extract the feature of the speech signal of the target object from the first speech features based on the processed mask features, so as to more accurately acquire the speech signal of the target object from the speech signal based on the more accurate feature of the speech signal of the target object.

[0129] It should be noted that the step 305 is an optional step, and the server can directly execute the step 306 after executing the step 304.

[0130] 306. The server extracts target speech features from the plurality of first speech features based on the mask features.

[0131] In the embodiment of the present application, the target speech features are features of the speech signal of the target object in the speech signal. The server splices the plurality of first speech features to obtain spliced features; and multiplies elements of the mask features and the spliced features to obtain the target speech features. The element multiplication refers to multiplying any element in the mask features and an element at a corresponding position in the spliced features. By multiplying the mask features and the spliced features, element values of speech features not belonging to the target object can be set to zero or a smaller value, and element values of speech features belonging to the target object can be set to a larger value, so as to obtain the target speech features.

[0132] 307. The server decodes the target speech features to obtain the speech signal of the target object.

[0133] In the embodiments of the present application, since the target speech feature is obtained through multiple first speech features, each first speech feature corresponds to a time resolution, and therefore the target speech feature corresponds to multiple time resolutions. The server decodes the target speech feature based on the multiple time resolutions in step 301 to obtain the speech signal of the target object. That is, the server uses a convolution kernel of the same size as in step 301 to perform inverse convolution on the target speech feature. Based on the inverse convolution operation, the server can restore the target speech feature to a time-domain waveform signal, i.e., the speech signal of the target object.

[0134] For example, continuing to refer to Figure 4 Or Figure 5 , the server inputs the target speech feature into a speech decoder, which includes a one-dimensional convolutional neural network with three convolution kernel sizes different from those of the speech encoder. Based on the convolutional neural network, the target speech feature can be inversely convolved to obtain the speech signal of the target object.

[0135] In order to more clearly describe the scheme in the embodiments of the present application, the following will be described from two aspects of single speech channel and multiple speech channels.

[0136] On the one hand, the server can obtain the target speech signal from the speech signal of the single speech channel. For example, continuing to refer to Figure 4 First, the server obtains the speech signal of the single speech channel. Then, the server inputs the speech signal into a speech encoder for extracting the speech features of the speech signal. The server extracts the features of the input speech signal through the speech encoder using a one-dimensional convolutional neural network with different convolution kernel sizes to obtain multiple first speech features. Then, the server weights the first speech features based on a channel attention mechanism to obtain the second speech features corresponding to the first speech features. In addition, the server also obtains the reference speech signal of the target object, inputs the reference speech signal into a speaker encoder, and obtains the speech features of the target object through the speaker encoder. Then, the server inputs the multiple second speech features and the speech features of the target object into a speaker extractor. The server obtains the mask features corresponding to the speech signal of the target object through the speaker extractor. Then, the server processes the sub-mask features in the mask features based on a context mechanism. Then, the server element-wise multiplies the processed mask features and the multiple first speech features of the speech signal to obtain the target speech feature. Finally, the server decodes the target speech feature through a speech decoder to obtain the target speech signal.

[0137] On the other hand, the server can obtain the target speech signal from the speech signal of the multiple speech channels. For example, continuing to refer to Figure 5First, the server acquires speech signals from multiple speech channels and inputs these signals into multiple speech encoders. Each speech channel corresponds to one speech encoder. For any given speech channel, the server obtains its speech features using the corresponding speech encoder. Then, the server inputs these speech features from multiple channels into a preprocessing module, which extracts spatial features from them to obtain spatial information features. Next, the server adds the speech features of the target speech channel to the spatial information features to obtain the first speech feature. The method by which the server obtains the target speech signal using the first speech feature is the same as the method for obtaining the target speech signal from a single speech channel, and will not be elaborated further here.

[0138] It should be noted that the solution proposed in the embodiments of this application can achieve significant results in tasks requiring the extraction of speech signals from a target object. For example, Figure 7 This is a schematic diagram of the spectrogram of a speech signal of a target object according to an embodiment of this application. See also Figure 7 , Figure 7 (a) in the example shows a spectrogram of a speech signal to be extracted. Figure 7 (b) in the example shows a spectrogram of a speech signal of a target object extracted using the SpEx+ model. Figure 7 (c) exemplarily illustrates a spectrogram of a speech signal of a target object extracted using the scheme provided in the embodiments of this application. Figure 7 (d) in the example shows a spectrogram of a speech signal from a real target object. Figure 7 The spectrogram shown in (b) and Figure 7 The spectrogram shown in (c) is respectively with Figure 7 Compare this with the spectrogram shown in (d) above. It can be seen that... Figure 7 The spectrogram shown in (b) is similar to... Figure 7 Compared to the spectrogram shown in (d), there is an oversuppression phenomenon, that is, the speech signal of the target object extracted by the SpEx+ model is suppressed in the dashed box part, and is not the complete speech signal of the target object. Figure 7 The spectrogram shown in (c) is the closest. Figure 8 The spectrogram shown in (d) is an example. This means that the speech signal of the target object extracted using the scheme provided in this application is more accurate than the speech signal of the target object extracted using the SpEx+ model.

[0139] The scheme provided in the embodiments of the present application can suppress the feature channels of irrelevant features in the speech feature, improve the performance of the speech feature, and further obtain more accurate speech signals of the target object from the speech signals.

[0140] Figure 8 is a structural schematic diagram of a speech signal acquisition device provided by the embodiments of the present application. Referring to Figure 9 The device comprises a first extraction module 801, a first processing module 802 and a first acquisition module 803.

[0141] The first extraction module 801 is configured to perform feature extraction on the speech signal according to multiple time resolutions to obtain multiple first speech features, the first speech feature corresponding to a time resolution, and the first speech feature comprising multiple feature channels.

[0142] The first processing module 802 is configured to, for any first speech feature, process the multiple feature channels in the first speech feature based on an attention mechanism to obtain a second speech feature corresponding to the first speech feature.

[0143] The first acquisition module 803 is configured to acquire the speech signal of the target object from the speech signal based on the multiple first speech features, the multiple second speech features and the speech feature of the target object.

[0144] In some embodiments, Figure 9 is a structural schematic diagram of another speech signal acquisition device provided by the embodiments of the present application. Referring to Figure 10 The first extraction module 801 comprises:

[0145] The first extraction unit 901 is configured to, for any speech channel, perform feature extraction on the signal of the speech channel according to multiple time resolutions to obtain multiple speech features of the speech channel.

[0146] The second extraction unit 902 is configured to perform spatial feature extraction on the speech features of the multiple speech channels to obtain spatial information features, the spatial information features being used to represent the spatial relationship between the speech features of the multiple speech channels.

[0147] The summation unit 903 is configured to add the multiple speech features of the target speech channel and the spatial information features respectively to obtain multiple first speech features, the target speech channel being the channel corresponding to the target object.

[0148] In some embodiments, the first acquisition module 803 comprises:

[0149] The determination unit 904 is configured to determine a mask feature based on the multiple second speech features and the speech feature of the target object.

[0150] The third extraction unit 905 is configured to extract target speech features from the plurality of first speech features based on the mask features.

[0151] The decoding unit 906 is configured to decode the target speech features to obtain a speech signal of the target object.

[0152] In some embodiments, the determination unit 904 is configured to normalize the plurality of second speech features respectively, and determine the mask features based on the normalized plurality of second speech features and the speech features of the target object.

[0153] In some embodiments, the third extraction unit 905 is configured to concatenate the plurality of first speech features to obtain concatenated features, and perform element multiplication on the mask features and the concatenated features to obtain the target speech features.

[0154] In some embodiments, the apparatus further includes:

[0155] The second acquisition module 804 is configured to acquire a reference speech signal, the reference speech signal being a speech signal of a target object.

[0156] The second extraction module 805 is configured to perform feature extraction on the reference speech signal according to a plurality of time resolutions to obtain a plurality of third speech features, the third speech features corresponding to the time resolutions.

[0157] The second processing module 806 is configured to perform feature embedding processing on the plurality of third speech features to obtain the speech features of the target object.

[0158] In some embodiments, the apparatus further includes:

[0159] The third acquisition module 807 is configured to, for any sub-mask feature, acquire at least two sub-mask features adjacent in time sequence to the sub-mask feature.

[0160] The fusion module 808 is configured to fuse the at least two sub-mask features and the sub-mask feature to obtain a fused sub-mask feature.

[0161] Embodiments of the present application provide an acquisition apparatus of a speech signal. By processing a plurality of feature channels in speech features based on an attention mechanism, irrelevant feature channels in the speech features can be suppressed, the performance of the speech features can be improved, and a more accurate speech signal of a target object can be acquired from the speech signal.

[0162] It should be noted that the voice signal acquisition device provided by the above embodiment is used to acquire voice signals, and the above-mentioned division of the functional modules is used for illustration. In actual application, the above-mentioned functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the voice signal acquisition device and the voice signal acquisition method provided by the above embodiment belong to the same concept, and the specific implementation process is described in the method embodiment, which will not be repeated here.

[0163] In the embodiments of the present application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can be used as an execution subject to implement the technical solutions provided by the embodiments of the present application. When the computer device is configured as a server, the server can be used as an execution subject to implement the technical solutions provided by the embodiments of the present application. The technical solutions provided by the present application can also be implemented through the interaction between the terminal and the server, and the embodiments of the present application do not limit this.

[0164] Figure 10 is a structural block diagram of a terminal 1000 according to an embodiment of the present application. The terminal 1000 can be a portable mobile terminal, such as a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a notebook computer, or a desktop computer. The terminal 1000 can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, and other names.

[0165] Generally, the terminal 1000 includes a processor 1001 and a memory 1002.

[0166] The processor 1001 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 1001 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 1001 can also include a main processor and a coprocessor, the main processor being a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit), and the coprocessor being a low-power processor for processing data in a standby state. In some embodiments, the processor 1001 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 1001 can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.

[0167] The memory 1002 can include one or more computer-readable storage media that can be non-transitory. The memory 1002 can also include high-speed random access memory and nonvolatile, computer-readable storage media such as one or more magnetic disk storage devices, flash memory devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one computer program for being executed by the processor 1001 to implement the method for obtaining a speech signal provided by the method embodiments in the present application.

[0168] In some embodiments, the terminal 1000 can also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002, and the peripheral device interface 1003 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 1003 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.

[0169] The peripheral interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002 and the peripheral interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002 and the peripheral interface 1003 can be implemented on a separate chip or circuit board, and the present embodiments are not limited in this regard.

[0170] The radio frequency circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1004 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. In some embodiments, the radio frequency circuit 1004 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 1004 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1004 can also include NFC (Near Field Communication) related circuitry, which is not limited by the present application.

[0171] The display screen 1005 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 1005 is a touch display screen, the display screen 1005 is further configured to capture touch signals on or above the surface of the display screen 1005. The touch signals can be input to the processor 1001 as control signals for processing. In this case, the display screen 1005 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 1005 can be one, disposed on the front panel of the terminal 1000; in other embodiments, the display screen 1005 can be at least two, respectively disposed on different surfaces of the terminal 1000 or in a folding design; in other embodiments, the display screen 1005 can be a flexible display screen, disposed on a curved surface or a folding surface of the terminal 1000. Even, the display screen 1005 can also be disposed in an irregular shape other than a rectangle, i.e., a special-shaped screen. The display screen 1005 can be made of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.

[0172] The camera assembly 1006 is configured to capture images or videos. In some embodiments, the camera assembly 1006 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, the rear camera is at least two, which is any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 1006 can further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0173] The audio circuit 1007 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into an electrical signal input to the processor 1001 for processing, or input to the radio frequency circuit 1004 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, respectively arranged at different parts of the terminal 1000. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signal from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker can be a traditional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, not only can it convert electrical signals into sound waves that humans can hear, but it can also convert electrical signals into sound waves that humans cannot hear for ranging purposes. In some embodiments, the audio circuit 1007 can also include a headphone jack.

[0174] The power supply 1008 is used to supply power to each component in the terminal 1000. The power supply 1008 can be alternating current, direct current, disposable battery or rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can be a wired charging battery or a wireless charging battery. The wired charging battery is a battery charged through a wired line, and the wireless charging battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0175] In some embodiments, the terminal 1000 also includes one or more sensors 1009. The one or more sensors 1009 include, but are not limited to, an acceleration sensor 1010, a gyroscope sensor 1011, a pressure sensor 1012, an optical sensor 1013, and a proximity sensor 1014.

[0176] The acceleration sensor 1010 can detect the acceleration in three coordinate axes of the coordinate system established by the terminal 1000. For example, the acceleration sensor 1010 can be used to detect the components of the gravitational acceleration in three coordinate axes. The processor 1001 can control the display screen 1005 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1010. The acceleration sensor 1010 can also be used for game or user motion data collection.

[0177] The gyroscope sensor 1011 can detect the body direction and rotation angle of the terminal 1000, and the gyroscope sensor 1011 can collect 3D actions of the user on the terminal 1000 in cooperation with the acceleration sensor 1010. The processor 1001 can realize the following functions according to the data collected by the gyroscope sensor 1011: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization when shooting, game control, and inertial navigation.

[0178] The pressure sensor 1012 can be disposed at the side frame of the terminal 1000 and / or the lower layer of the display screen 1005. When the pressure sensor 1012 is disposed at the side frame of the terminal 1000, the holding signal of the user to the terminal 1000 can be detected, and the left-hand or right-hand recognition or the shortcut operation can be performed by the processor 1001 according to the holding signal collected by the pressure sensor 1012. When the pressure sensor 1012 is disposed at the lower layer of the display screen 1005, the operable control on the UI interface can be controlled by the processor 1001 according to the pressure operation of the user to the display screen 1005. The operable control includes at least one of the button control, the scroll bar control, the icon control, and the menu control.

[0179] The optical sensor 1013 is used to collect the ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 according to the ambient light intensity collected by the optical sensor 1013. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameter of the camera assembly 1006 according to the ambient light intensity collected by the optical sensor 1013.

[0180] The proximity sensor 1014, also called the distance sensor, is usually disposed at the front panel of the terminal 1000. The proximity sensor 1014 is used to collect the distance between the user and the front of the terminal 1000. In one embodiment, when the proximity sensor 1014 detects that the distance between the user and the front of the terminal 1000 gradually decreases, the display screen 1005 is switched from the bright screen state to the off-screen state by the processor 1001; when the proximity sensor 1014 detects that the distance between the user and the front of the terminal 1000 gradually increases, the display screen 1005 is switched from the off-screen state to the bright screen state by the processor 1001.

[0181] Those skilled in the art can understand that the structures shown in the above embodiments are not a limitation on the terminal 1000, and the terminal 1000 can include more or less components than those shown in the figures, or combine certain components, or adopt different component arrangements. Figure 11

[0182] ​ ​A structural schematic diagram of a server is provided according to an embodiment of the present application. The server 1100 can have great differences due to different configurations or performances, and can include one or more processors (Central Processing Units, CPUs) 1101 and one or more memories 1102. The memory 1102 stores at least one computer program, which is loaded and executed by the processor 1101 to implement the method for acquiring a voice signal provided by each of the above-mentioned method embodiments. Of course, the server can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for implementing device functions, and the like, so as to perform input and output. The server can also include other components for implementing device functions, which are not described herein.

[0183] The embodiment of the present application further provides a computer readable storage medium, which stores at least one computer program. The at least one computer program is loaded and executed by a processor of a computer device to implement operations performed by the computer device in the method for acquiring a voice signal provided by the above-mentioned embodiments. For example, the computer readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0184] The embodiment of the present application further provides a computer program product, which includes computer program code stored in a computer readable storage medium. A processor of a computer device reads the computer program code from the computer readable storage medium, and executes the computer program code, so that the computer device performs the method for acquiring a voice signal provided in the various optional implementation manners.

[0185] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a Read-Only Memory, a magnetic disk or an optical disk, and the like.

[0186] The above-mentioned is an optional embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, and the like made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method of acquiring a voice signal, characterized by, The method comprises: characteristic extraction on the speech signal according to a plurality of time resolutions to obtain a plurality of first speech characteristics, the first speech characteristics corresponding to the time resolutions, the first speech characteristics comprising a plurality of feature channels; for any first speech characteristic, processing a plurality of feature channels in the first speech characteristic based on an attention mechanism to obtain a second speech characteristic corresponding to the first speech characteristic; based on the plurality of second speech characteristics and the speech characteristic of the target object, determining a mask characteristic, the mask characteristic comprising a plurality of sub-mask characteristics; for any sub-mask characteristic in the mask characteristic, obtaining at least two sub-mask characteristics sequentially adjacent to the sub-mask characteristic, and fusing the at least two sub-mask characteristics and the sub-mask characteristic to obtain a fused sub-mask characteristic; based on the processed mask characteristic, extracting a target speech characteristic from the plurality of first speech characteristics; decoding the target speech characteristic to obtain the speech signal of the target object.

2. The method of claim 1, wherein, The speech signal comprises signals of a plurality of speech channels; the characteristic extraction on the speech signal according to a plurality of time resolutions to obtain a plurality of first speech characteristics comprises: for any speech channel, characteristic extraction on the signal of the speech channel according to a plurality of time resolutions to obtain a plurality of speech characteristics of the speech channel; spatial feature extraction on the speech characteristics of the plurality of speech channels to obtain spatial information characteristics, the spatial information characteristics being used to represent the spatial relationship between the speech characteristics of the plurality of speech channels; adding the plurality of speech characteristics of a target speech channel and the spatial information characteristics respectively to obtain the plurality of first speech characteristics, the target speech channel being the channel corresponding to the target object.

3. The method of claim 1, wherein, The determination of the mask characteristic based on the plurality of second speech characteristics and the speech characteristic of the target object comprises: normalizing the plurality of second speech characteristics respectively; determining the mask characteristic based on the normalized plurality of second speech characteristics and the speech characteristic of the target object.

4. The method of claim 1, wherein, The extraction of the target speech characteristic from the plurality of first speech characteristics based on the processed mask characteristic comprises: splicing the plurality of first speech characteristics to obtain spliced characteristics; element multiplication of the processed mask characteristic and the spliced characteristics to obtain the target speech characteristic.

5. The method of claim 1, wherein, The method further comprises: obtaining a reference speech signal, the reference speech signal being the speech signal of the target object; characteristic extraction on the reference speech signal according to the plurality of time resolutions to obtain a plurality of third speech characteristics, the third speech characteristics corresponding to the time resolutions; feature embedding processing on the plurality of third speech characteristics to obtain the speech characteristic of the target object.

6. An apparatus for acquiring a voice signal, characterized by comprising: The device comprises: a first extraction module for characteristic extraction on a speech signal according to a plurality of time resolutions to obtain a plurality of first speech characteristics, the first speech characteristics corresponding to the time resolutions, the first speech characteristics comprising a plurality of feature channels; The first processing module is configured to, for any first speech feature, process multiple feature channels in the first speech feature based on an attention mechanism to obtain a second speech feature corresponding to the first speech feature; The first obtaining module is configured to determine a mask feature based on the multiple second speech features and a speech feature of a target object, the mask feature including multiple sub-mask features; The third obtaining module is configured to, for any sub-mask feature in the mask feature, obtain at least two sub-mask features that are time-sequentially adjacent to the sub-mask feature; The fusion module is configured to fuse the at least two sub-mask features and the sub-mask feature to obtain a fused sub-mask feature. The first obtaining module is further configured to extract a target speech feature from the multiple first speech features based on the processed mask feature. The first obtaining module is further configured to decode the target speech feature to obtain a speech signal of the target object.

7. The apparatus of claim 6, wherein The speech signal includes signals of multiple speech channels; The first extracting module includes: The first extracting unit is configured to, for any speech channel, perform feature extraction on signals of the speech channel according to multiple time resolutions to obtain multiple speech features of the speech channel; The second extracting unit is configured to perform spatial feature extraction on the speech features of the multiple speech channels to obtain spatial information features, the spatial information features being used to represent spatial relationships between the speech features of the multiple speech channels; The summing unit is configured to add the multiple speech features of a target speech channel and the spatial information features respectively to obtain the multiple first speech features, the target speech channel being a channel corresponding to the target object.

8. The apparatus of claim 6, wherein, The first obtaining module includes a determining unit configured to: perform normalization on the multiple second speech features respectively; determine the mask feature based on the normalized multiple second speech features and the speech feature of the target object.

9. The apparatus of claim 6, wherein, The first obtaining module includes a third extracting unit configured to: splice the multiple first speech features to obtain spliced features; perform element multiplication on the processed mask feature and the spliced features to obtain the target speech feature.

10. The apparatus of claim 6, wherein The apparatus further includes: The second obtaining module is configured to obtain a reference speech signal, the reference speech signal being a speech signal of the target object; The second extracting module is configured to perform feature extraction on the reference speech signal according to the multiple time resolutions to obtain multiple third speech features, the third speech features corresponding to the time resolutions; The second processing module is configured to perform feature embedding processing on the multiple third speech features to obtain the speech feature of the target object.

11. A computer device, comprising: The computer device includes a processor and a memory, the memory being used to store at least one piece of computer program, the at least one piece of computer program being loaded and executed by the processor to perform the speech signal obtaining method in any one of claims 1 to 5.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of computer program, the at least one piece of computer program being used to perform the speech signal obtaining method in any one of claims 1 to 5.

13. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method for obtaining a speech signal as claimed in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Speaker identification method and system

    CN113611314A