Regional pickup method and related equipment
By embedding DOA information in the TSE method and filtering the voice data, the problem of insufficient robustness of the existing TSE method in complex environments is solved, and more efficient and accurate target speaker voice extraction is achieved.
Patent Information
- Application Number
- CN202510487729.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-27
AI Technical Summary
The existing TSE method is not robust in DOA mismatch and complex noise environments, making it difficult to effectively extract the voice signal of the target speaker.
By obtaining mixed voice data, arrival direction DOA information and beam width, DOA information is encoded and embedded, the beam width range is determined, the voice data is filtered, the noise reduction voice data is obtained, and voice extraction is performed to obtain the voice information of the target speaker in the target area.
In noisy environments, enhance the perception and spatial selectivity of the target speaker's direction, achieve dynamic locking of the target area, improve the clarity and accuracy of speech extraction, and enhance the utilization of space-time information of multi-microphone arrays.
Smart Images

Figure CN120220709A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of sound processing, and in particular, to a method for regional sound pickup and related devices. Background Art
[0002] Regional sound pickup is a technology for extracting sound signals from a specific spatial area, which is widely used in fields such as environmental sound collection, noise suppression, and target speaker extraction. Target Speaker Extraction (TSE) aims to extract the speech signal of a specific target speaker from a mixed speech signal containing multiple speakers, and has broad application prospects in fields such as remote conferencing and in-vehicle voice interaction.
[0003] Existing TSE methods usually rely on two main clues: speaker-specific voiceprint features and Direction of Arrival (DOA) features. Voiceprint features require pre-registered speech of the target speaker, which is often difficult to achieve in practical applications. In contrast, DOA features utilize the spatial information of the microphone array and can achieve target speaker extraction without relying on pre-registered speech. However, existing DOA feature-based TSE methods still face some challenges in practical applications, especially the lack of robustness in the case of DOA mismatch and complex noise environments. Therefore, there is an urgent need for a new TSE method that can make full use of the spatio-temporal information of the multi-microphone array and maintain high extraction performance in the case of DOA mismatch. Summary of the Invention
[0004] Embodiments of this application provide a method for regional sound pickup and related devices, which can solve the above problems. The technical solutions are as follows:
[0005] In a first aspect, embodiments of this application provide a method for regional sound pickup, the method comprising:
[0006] Obtaining mixed speech data, Direction of Arrival (DOA) information, and beam width;
[0007] Encoding the DOA information to obtain DOA vector data, and embedding the DOA vector data into the mixed speech data to obtain speech data to be processed;
[0008] Determining a beam width range based on the beam width, and filtering the speech data to be processed except for the speech data corresponding to the beam width range to obtain noise-reduced speech data;
[0009] Performing speech extraction on the noise-reduced speech data to obtain the speech information of the target speaker in the target area.
[0010] Second aspect, embodiments of the present application provide a regional sound pickup device, the device comprising:
[0011] An information acquisition module, configured to acquire mixed voice data, direction of arrival (DOA) information, and beam width;
[0012] A vector embedding module, configured to encode the DOA information to obtain DOA vector data, and embed the DOA vector data into the mixed voice data to obtain voice data to be processed;
[0013] A data filtering module, configured to determine a beam width range based on the beam width, and filter out voice data other than the voice data corresponding to the beam width range in the voice data to be processed, to obtain noise-reduced voice data;
[0014] A voice extraction module, configured to perform voice extraction on the noise-reduced voice data to obtain voice information of a target speaker in a target area.
[0015] Third aspect, embodiments of the present application provide a computer storage medium storing multiple instructions, which are adapted to be loaded and executed by a processor to perform the above method steps.
[0016] Fourth aspect, embodiments of the present application provide an electronic device, which may include: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the above method steps.
[0017] The beneficial effects brought by the technical solutions provided by some embodiments of the present application at least include:
[0018] In the present application, for the collected mixed voice data, the DOA information is encoded and embedded into the mixed voice data to obtain voice data to be processed. By embedding the DOA information, the perception of the direction of the target speaker and the spatial selectivity are enhanced, which helps to focus on the voice information in the target area in a noisy environment and achieve dynamic locking of the target area. Further, the beam width range is dynamically adjusted according to the specific value of the beam width, and voice data other than the voice data corresponding to the beam width range in the voice data to be processed is filtered to obtain noise-reduced voice data. The size of the required beam width range can be adjusted according to the actual application scenario, so as to adjust the spatial range of voice extraction, perform precise noise reduction on the voice data to be processed, improve the clarity and accuracy of extracting voice information from the noise-reduced voice data, and provide strong support for subsequent tasks such as speech recognition. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic structural diagram of a regional sound pickup method provided by an embodiment of the present application;
[0021] Figure 2 It is a schematic scenario diagram of a regional sound pickup method provided by an embodiment of the present application;
[0022] Figure 3 It is a schematic flowchart of a regional sound pickup method provided by an embodiment of the present application;
[0023] Figure 4 It is a schematic flowchart of a regional sound pickup method provided by an embodiment of the present application;
[0024] Figure 5 It is a schematic flowchart of a process for obtaining noise-reduced voice data provided by an embodiment of the present application;
[0025] Figure 6 It is a schematic flowchart of a model training process provided by an embodiment of the present application;
[0026] Figure 7 It is a schematic flowchart of a regional sound pickup method provided by an embodiment of the present application;
[0027] Figure 8 It is a schematic flowchart of a regional sound pickup method provided by an embodiment of the present application;
[0028] Figure 9 It is a schematic structural diagram of a broadband layer and a narrowband layer provided by an embodiment of the present application;
[0029] Figure 10 It is a schematic structural diagram of a regional sound pickup device provided by an embodiment of the present application;
[0030] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0032] In the description of the present application, it should be understood that terms such as "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In the description of the present application, it should be noted that unless otherwise clearly specified and limited, "including" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise stated, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0033] The present application will be described in detail below with reference to specific embodiments.
[0034] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the features, information, and data involved in the present application are all obtained under full authorization.
[0035] As Figure 1 shown, Figure 1 is a schematic flowchart of a regional sound pickup method provided by an embodiment of the present application. Figure 1 It at least includes a server 101 that executes the regional sound pickup method, and further includes a plurality of electronic devices for uploading mixed voice data. The plurality of electronic devices at least include an electronic device 1021, an electronic device 1022, and an electronic device 1023. It can be understood that Figure 1 the numbers of the server and the electronic devices shown in
[0036] The above-mentioned server 101 can be a single server device, such as: a rack-mounted, blade, tower, or cabinet-style server device, or a hardware device with strong computing capabilities such as a workstation or a mainframe computer; it can also be a server cluster composed of multiple servers. The servers in the service cluster can be composed symmetrically, where each server is functionally equivalent and has an equivalent status in the transaction link, and each server can provide services externally independently. Providing services independently can be understood as not requiring the assistance of another server.
[0037] For example, the server is multiple physical servers, and the multiple physical servers are independent in terms of hardware. Or, the server is multiple virtual servers, and the multiple virtual servers are deployed in the same hardware resource pool. The deployment methods of virtual servers include, but are not limited to: VMware, Virtual Box, and Virtual PC.
[0038] It can be understood that the server 101 also has other service capabilities and functions to complete the tasks in the following embodiments. For example, the server 101 also provides portal services, resource management services, and CI / CD services, etc.
[0039] The electronic device includes, but is not limited to: wearable devices, handheld devices, personal computers, tablet computers, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem, etc. In different networks, the electronic device can be called by different names, such as: user equipment, access terminal, user unit, user station, mobile station, mobile device, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic device in a 5G network or future evolved network, etc.
[0040] In the embodiments of the present application, a sound collection device is provided on each of the electronic devices such as the electronic device 1021, the electronic device 1022, and the electronic device 1023. The sound collection device is used to collect voice data of at least one speaker in a certain scenario. For example, the sound collection device can be a condenser microphone, a dynamic microphone, a moving iron microphone, or an electret microphone, etc. It can be understood that devices such as a sound card, an audio processing chip, and a sound output device for converting an audio signal into a digital signal are provided on the electronic device to cooperate with the sound collection device to process the collected initial voice data.
[0041] As Figure 2 shown, Figure 2It is a schematic diagram of the scenario of a regional sound pickup method provided by an embodiment of the present application. In this regional sound pickup scenario, there are speaker 201, speaker 202, speaker 203, speaker 204, and speaker 205. The above-mentioned multiple speakers will take turns speaking or speak simultaneously. The electronic device will collect the initial voice data in this scenario and send the initial voice data to the server 101 so that the server 101 can execute the regional sound pickup method based on the initial voice data.
[0042] Multiple electronic devices and the server can communicate through a communication link established by a communication protocol. For example: Among them, the network can be a wireless network or a wired network. The wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network, or a Bluetooth network. The wired network includes but is not limited to an Ethernet, a universal serial bus (USB), or a controller area network. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data (such as the target compressed package) exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0043] After research by the inventor, it is found that there are still many problems in the existing TSE method. First, the inaccuracy of DOA feature measurement (such as the calibration error of the microphone array and environmental noise) will significantly affect the performance of TSE. Second, when the existing TSE method is given DOA features, it usually faces the problem of DOA feature mismatch, resulting in uncertain output. This ambiguity may cause interference signals from other directions to be mixed into the extracted voice signal. Especially in a multi-speaker scenario, when the directions of the target speaker and the interfering speaker are close, the extraction effect will be significantly reduced. In addition, the existing methods do not fully utilize the spatio-temporal information collected by multiple microphone arrays. For example, although traditional beamforming methods (such as MVDR) can suppress noise and interference to a certain extent, their performance is still limited in a complex noise environment. Therefore, the present application proposes a regional sound pickup method to solve this problem.
[0044] In one embodiment, as Figure 3 shown, it is a schematic flowchart of a regional sound pickup method provided by an embodiment of the present application. This method can be implemented depending on a computer program and can run on a regional sound pickup device based on the von Neumann architecture. This computer program can be integrated into an application or run as an independent tool class application.
[0045] Specifically, the regional sound pickup method includes:
[0046] S101. Obtain mixed speech data, direction of arrival (DOA) information, and beamwidth.
[0047] The mixed speech data can be understood as speech data containing background noise or multiple sound sources (i.e., multiple speakers). In this embodiment, for the noisy multi-channel speech data collected by a sound collection device including multiple microphones, operations such as frame segmentation, windowing, and short-time Fourier transform are performed to obtain the mixed speech data.
[0048] As Figure 4 shown, Figure 4 it is a schematic flowchart of a regional sound pickup method provided by an embodiment of the present application. The multi-channel mixed speech data 301 is collected by an X ∈ R M×L microphone array, where L is the number of time samples. Frame segmentation, windowing processing, and short-time Fourier transform (STFT) are performed on it to convert it into a spectrogram with dimensions of 2M × T × f. Among them, M is the number of microphones, and T and F respectively represent the number of time and frequency bins. In this way, the multi-channel mixed speech data as a time-domain signal can be converted into a frequency-domain signal, facilitating subsequent analysis and processing of the frequency characteristics of the signal.
[0049] Furthermore, the spectrogram obtained from the processed multi-channel mixed speech data is processed using a convolutional input layer to expand the channel dimension to C, so as to obtain the mixed speech data. In this process, the spatial characteristics and local time-frequency information corresponding to the mixed speech data are encoded, providing a feature sequence for the subsequent processing of the regional sound pickup deep neural network.
[0050] The beamwidth is an important parameter used to describe the radiation directivity of an antenna or an acoustic device. A narrow beamwidth means that the electronic device has a strong directivity in collecting speech data, can accurately focus on a certain area, and reduce noise interference from other areas. A wide beamwidth means that the speech data reception area of the electronic device is wider, can cover a larger area, but has a weaker directivity and may receive more interference or noise.
[0051] In this application, the acquisition of the beam width can be based on the fixed parameters of the electronic device. For example, the parameters of the sound collection device set on the electronic device are obtained. For instance, the layout of the microphone array (such as a linear array, two-dimensional array) determines the width of its beam width. The larger the size of the microphone array, the smaller the beam width and the narrower the range that can be focused; the smaller the size of the microphone array, the larger the beam width and the wider the directions that can be received.
[0052] In another embodiment, the beam width can be determined according to the sound collection scenario, and the beam width can be dynamically adjusted. For example, when the collection scenario requires high-precision sound collection, the value of the beam width is decreased. For another example, when the collection scenario requires covering a larger range of sound collection, the value of the beam width is increased.
[0053] The acquisition of DOA information can be calculated based on the multi-channel mixed speech data 301. For example, beamforming calculations are performed on the multi-channel mixed speech data 301 through the Steered Response Power-SRP method or the Minimum Variance Distortionless Response (MVDR) method to determine the DOA information. The multi-channel mixed speech data is obtained by performing processing such as noise reduction, framing, windowing, and short-time Fourier transform on the mixed speech, reducing the environmental noise existing in the multi-channel mixed speech data. Further, the DOA information is extracted from the multi-channel mixed speech data through methods such as the Minimum Variance Distortionless Response, which can improve the accuracy of the finally obtained DOA information and avoid the influence of inaccurate DOA information in the prior art on the process of regional sound pickup.
[0054] S102: Encode the DOA information to obtain DOA vector data, and embed the DOA vector data into the mixed speech data to obtain the speech data to be processed.
[0055] The purpose of the encoding process is to convert the DOA information into a form suitable for embedding into the mixed speech data. The encoding method can be simple binary encoding or a certain compression encoding technique, such that the data volume of the DOA vector is small and can be transmitted quickly.
[0056] In the regional sound pickup deep neural network 304, the DOA vector data is broadcast along the time dimension and applied to the mixed speech data through element-wise multiplication to obtain the speech data to be processed. Broadcasting can be understood as replicating and expanding a smaller DOA vector to be the same size as the mixed speech signal in the time dimension. Specifically, the DOA vector is extended to be the same as the entire time series based on the estimated angle at each moment. Further, by performing element-wise multiplication on the DOA vector and the mixed speech data, that is, by weighting, the different sound sources in the mixed speech data are enhanced or suppressed to obtain the speech data to be processed, thereby strengthening the selection of the target area and the target sound source (i.e., the target speaker).
[0057] S103. Determine a beam width range based on the beam width, and filter the speech data to be processed except for the speech data corresponding to the beam width range to obtain noise-reduced speech data.
[0058] The beam width range is determined based on the beam width, and a range is set with the value of the beam width as the center. The difference between the upper and lower boundary values corresponding to the beam width range is not exactly the same for different values of the beam width. For example, when the beam width is 30°, the beam width range is 15° - 45°, and the difference corresponding to this beam width range is 30°. When the beam width is 90°, the beam width range is 60° - 120°, and the difference corresponding to this beam width range is 60°. This is because the signal acquisition accuracy requirements corresponding to different scenarios are different. For high-precision acquisition, the corresponding beam width is smaller and the difference corresponding to the beam width range is smaller. For a wider acquisition range, the corresponding beam width is larger and the difference corresponding to the beam width range is larger.
[0059] The speech data within the beam width range is considered to be the speech data collected from the target area in the entire scene. In the regional sound pickup deep neural network 304, the target area is determined based on the DOA information embedded in the speech data to be processed, and further a spatial filter is determined through the beam width range corresponding to the beam width 303. This spatial filter can enhance the signal of the target speaker in the target area within this beam width range while suppressing the noise in other directions. This filter is actually a weighting process.
[0060] For example, when the beam width is 30° and the beam width range is 15° - 45°, only the speech data corresponding to 15° - 45° in the speech data to be processed will be enhanced, and the speech data obtained from other angles will be filtered, thereby obtaining noise-reduced speech data with reduced sound from other areas.
[0061] For the collected mixed speech data, the DOA information representing the spatial position characteristics is encoded and embedded into the mixed speech data to obtain the speech data to be processed, enhancing the perception of the target speaker's direction and spatial selectivity. In addition, by fully utilizing the spatio-temporal information collected by the multi-microphone array, the speech data in the speech data to be processed other than the speech data corresponding to the beam width range is filtered through the beam width, adjusting the spatial range of speech extraction, and further locking the target speaker's direction, which can improve the matching degree of DOA information in the sound pickup processing and solve the problem of DOA information mismatch in the prior art.
[0062] S104. Perform speech extraction on the noise-reduced speech data to obtain the speech information of the target speaker in the target area.
[0063] Based on the noise-reduced speech data, the regional sound pickup deep neural network 304 performs a speech extraction task to separate the speech information of the target speaker in the target area from the noise-reduced speech data. For example, the target speaker is Figure 2 the speaker 201 shown in
[0064] The regional sound pickup deep neural network 304 can be a deep learning network such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a variational autoencoder (VAE), including, for example, the deep speech separation network (DeepClustering) or the time-frequency masking network (Time-Frequency Masking Network) in deep learning. Combining the noise-reduced speech data obtained through the beam width range and the DOA vector embedded in the noise-reduced speech data, it learns to separate the speech information of the target speaker in the target area.
[0065] Finally, the regional sound pickup deep neural network 304 maps the processed speech information to the estimated target STFT coefficients through a linear output layer, and then applies the inverse short-time Fourier transform (iSTFT) to convert the speech information as a frequency-domain signal back to the time domain, reconstructing the enhanced and extracted speech information of the target speaker, and completing the task of extracting the speech information of the target speaker in the target area of the multi-channel speech data 301.
[0066] In the present application, for the collected mixed speech data, the DOA information is encoded and embedded into the mixed speech data to obtain the speech data to be processed. By embedding the DOA information, the perception of the direction of the target speaker and the spatial selectivity are enhanced, which helps to focus on the speech information in the target area in a noisy environment and achieve dynamic locking of the target area. Further, according to the specific value of the beam width, the beam width range is dynamically adjusted to filter the speech data in the speech data to be processed except for the speech data corresponding to the beam width range, so as to obtain the noise-reduced speech data. The size of the required beam width range can be adjusted according to the actual application scenario, thereby adjusting the spatial range of speech extraction, accurately reducing the noise of the speech data to be processed, and improving the clarity and accuracy of extracting speech information from the noise-reduced speech data, providing strong support for subsequent tasks such as speech recognition.
[0067] Based on the embodiments such as Figure 1 - Figure 4 shown, in one embodiment, S102 specifically includes: performing low-dimensional encoding on the DOA information through circular positional encoding to obtain DOA vector data, and embedding the DOA vector data into the mixed speech data to obtain the speech data to be processed.
[0068] Circular Positional Encoding is an encoding method used to process data with periodic or circular structures, and specifically considers the periodic characteristics of the DOA information to be processed.
[0069] The DOA information representing the source direction of the mixed speech data includes angle data, so the DOA information is a typical periodic data, and 0° and 360° represent the same direction. In this embodiment, circular positional encoding is used to process the DOA information, embed the DOA information into the periodic space, generate low-dimensional and continuous DOA vector data, and the DOA vector data is broadcast along the time dimension.
[0070] The prior art usually uses high-dimensional one-hot encoding to represent the DOA information, resulting in low computational efficiency and inability to effectively capture the continuity of the direction. This embodiment uses circular positional encoding to perform low-dimensional embedding of the DOA information, improving the selectivity for the target space and the ability to extract speech information.
[0071] Based on the embodiments such as Figure 1 - Figure 4 shown, in one embodiment, S102 specifically includes: encoding the DOA information to obtain DOA vector data, and embedding the DOA vector data into the mixed speech data through clue encoding to obtain the speech data to be processed.
[0072] Cue Encoding can be understood as a technique for embedding additional information (such as direction information, location information, etc.) into mixed speech data that serves as the main data, with the aim of enabling a regional voice pickup deep learning network to better utilize this additional information when processing the speech data to be processed. For example, it can better determine the target area where the sound source is located, so as to separate and process the sound source from the speech data to be processed.
[0073] Embedding the DOA vector into the mixed speech data to obtain the speech data to be processed can be achieved by directly splicing the DOA vector as an additional feature onto the feature vector of the mixed speech data. For example, after extracting the feature vector of the mixed speech data in the frequency domain or time domain, the DOA vector can be spliced as an additional vector at the end of the feature vector of the mixed speech data to obtain the feature vector of the speech data to be processed. When the regional voice pickup deep learning network processes the speech data to be processed, it processes it in combination with the DOA vector as context information.
[0074] Embedding the DOA vector into the mixed speech data to obtain the speech data to be processed can also be done by calculating attention weights based on the DOA vector. Usually, through a neural network or other learning methods, the DOA vector is mapped to a weight vector, and through the attention mechanism, the regional voice pickup deep learning network is guided to assign different degrees of attention to the speech data corresponding to each angle in the mixed speech data in combination with the DOA vector. This method can dynamically adjust the influence of the DOA information according to the specific task, so as to better adapt to complex scenarios.
[0075] In this embodiment, the DOA vector obtained by encoding the DOA information is embedded into the mixed speech data through a cue encoder to obtain the speech data to be processed including the DOA information as a cue, enhancing the perception of the target speaker's direction and spatial selectivity of the regional voice pickup deep network, and helping to focus on the speech signal in the target area in a complex environment.
[0076] Based on the embodiments as shown in Figure 1 - Figure 4 Please also refer to the embodiments as shown in Figure 5 the embodiments as shown. Figure 5 FIG. is a schematic flowchart of a process for obtaining noise-reduced speech data provided by an embodiment of the present application. S103 specifically includes the following steps:
[0077] S201. Based on the association relationship between the preset beam width and the beam width range, determine the beam width range associated with the beam width based on the beam width.
[0078] The beam width range is determined based on the beam width, and a range is set with the value of the beam width as the center. When the value of the beam width is different, the difference between the upper bound value and the lower bound value corresponding to the beam width range is not exactly the same. The correlation between the beam width and the beam width range includes a positive correlation between the beam width and the difference corresponding to the beam width range.
[0079] For example, when the beam width is 30°, the beam width range is 15° - 45°, and the difference corresponding to this beam width range is 30°. When the beam width is 90°, the beam width range is 60° - 120°, and the difference corresponding to this beam width range is 60°. This is because the requirements for signal acquisition accuracy in different scenarios are different. For high-precision acquisition, the corresponding beam width is smaller and the difference corresponding to the beam width range is smaller. For a wider acquisition range, the corresponding beam width is larger and the difference corresponding to the beam width range is larger.
[0080] In one embodiment, before obtaining the mixed voice data, the direction of arrival (DOA) information, and the beam width, it further includes: obtaining a beam width matching the pick-up scenario according to the pick-up scenario.
[0081] The pick-up scenario can be understood as the acquisition environment when multi-channel voice data is acquired by a sound acquisition device, and the accuracy requirement for extracting voice information from the multi-channel voice data. When the acquisition environment requires a small sound acquisition angle range and a high accuracy requirement for extracting voice information, the beam width matching this pick-up scenario is smaller, and the difference between the upper bound value and the lower bound value corresponding to this beam width range is smaller. When the acquisition environment requires a wider sound acquisition angle coverage, the beam width matching this pick-up scenario is larger, and the difference between the upper bound value and the lower bound value corresponding to this beam width range is larger.
[0082] In this embodiment, by setting different beam widths for different pick-up scenarios and dynamically adjusting the beam width range for different beam widths, the accuracy of the beam width can be improved, so as to fully utilize the beam width to filter the voice data in the to-be-processed voice data except for the voice data corresponding to the beam width range, and improve the noise reduction effect on the to-be-processed voice data.
[0083] S202: Encode the beam width range to obtain a range mask, and filter the voice data in the to-be-processed voice data except for the voice data corresponding to the beam width range according to the range mask to obtain noise-reduced voice data.
[0084] Specifically, the beam width information is encoded by one-hot encoding, and a range mask is generated through a linear layer and 1×1 Conv2D. This range mask can dynamically adjust the expected beam width within the beam width range, thereby determining the speech data other than the speech data corresponding to the expected beam width in the speech data to be processed as noise, and filtering out the above-mentioned noise to obtain noise-reduced speech data.
[0085] In one embodiment, the beam width range is encoded to obtain a range mask, and the speech data other than the speech data corresponding to the beam width range in the speech data to be processed is filtered according to the range mask to obtain the speech data to be processed for noise reduction; the speech data to be processed for noise reduction and the speech data to be processed are subjected to residual connection processing to obtain noise-reduced speech data.
[0086] Residual Connection, also known as Skip Connection or Jump Link, is a technique in neural network structures. Its core idea is to directly "jump" connect the input data with the output of the network, avoiding the gradual loss of information in a multi-layer network structure. This connection method enables the network to directly pass the input signal to the subsequent layers, helping to alleviate the vanishing gradient problem, thereby improving the training effect and accelerating convergence.
[0087] In this embodiment, after the speech data to be processed is denoised to obtain the speech data to be processed for noise reduction, the speech data to be processed for noise reduction and the speech data to be processed are subjected to residual connection processing. The residual connection is used to retain the key information in the speech data to be processed as the original signal, ensuring that while suppressing the noise in the speech data to be processed, the important features of the speech information of the target speaker in the target area are not lost, and further improving the focusing ability for the target area and the target speaker.
[0088] Based on the embodiments as Figure 1 - Figure 4 shown, please also refer to the embodiments as Figure 6 shown. As Figure 6 shown, Figure 6 is a schematic flowchart of a model training provided by an embodiment of the present application. Before S101, it further includes:
[0089] S301. Obtain the pre-trained beam width and pre-trained mixed speech data, as well as the speech information of the speaker in the pre-trained mixed data.
[0090] First, through the DOA information from the active speech source and the mixed speech data, the regional sound pickup depth neural network is trained to initially have the ability to extract the speech information of the target speaker from the mixed speech data.
[0091] Further, obtain the pre-trained beam width and pre-trained mixed speech data, as well as the speech information of the speaker in the pre-trained mixed data. The obtaining method can be to obtain from a preset database, and this application does not impose any restrictions on this.
[0092] S302. Assign an initial beam width range to the pre-trained beam width, and extract initial speech information from the pre-trained mixed speech data based on the initial beam width range.
[0093] The second stage of the training of the regional voice pickup deep neural network is the beam adaptation stage. The goal is to enable the regional voice pickup deep neural network to have the ability to determine the most appropriate beam width range based on the input beam width, and to extract the speech information of the target speaker in the target area from the mixed speech data according to the beam width range.
[0094] In the training of this stage, randomly assign an initial beam width range to any pre-trained beam width, and extract initial speech information from the pre-trained mixed speech data based on the initial beam width range. For example, when the pre-trained beam width is 30°, configure the initial beam width range to be 25° - 35°, and extract initial speech information from the pre-trained mixed speech data based on this initial beam width range. The extraction process refers to the above S101 - S104 and will not be elaborated here.
[0095] S303. According to the initial speech information and the speech information, adjust the initial beam width range assigned to the pre-trained beam width until the speech information extracted from the pre-trained mixed speech data based on the beam width range corresponding to the pre-trained beam width meets the training conditions, then determine the beam width range corresponding to the pre-trained beam width, and construct the association relationship between the beam width and the beam width range.
[0096] Construct a loss function according to the initial speech information extracted during the training process and the speech information of the speaker in the pre-trained mixed data, and adjust the initial beam width range assigned to the pre-trained beam width according to the loss function. Train the regional voice pickup deep neural network multiple times until the speech information extracted from the pre-trained mixed speech data based on the beam width range corresponding to the pre-trained beam width meets the training conditions, then determine the beam width range corresponding to the pre-trained beam width.
[0097] Train the regional voice pickup deep neural network with multiple beam widths of different values to construct the association relationship between the beam width and the beam width range. During the training process, for the case where there is no voice source in the target area, the regional voice pickup deep neural network can generate a very small fixed signal to avoid unstable training.
[0098] In this embodiment, the regional voice pickup deep neural network is trained so that the trained regional voice pickup deep neural network can dynamically adjust the beam width range based on the beam width, thereby dynamically adjusting the spatial range of voice extraction, performing precise noise reduction on the voice data to be processed, and flexibly meeting the requirements of different scenarios.
[0099] In one embodiment, as Figure 7 shown, it is a schematic flowchart of a regional voice pickup method provided by an embodiment of the present application. This method can be implemented depending on a computer program and can run on a regional voice pickup device based on the von Neumann architecture. This computer program can be integrated in an application or run as an independent tool class application.
[0100] Specifically, the regional voice pickup method includes:
[0101] S401. Obtain mixed voice data, arrival DOA information, and beam width.
[0102] Referring to the above S101, it will not be elaborated here.
[0103] S402. Perform low-dimensional encoding on the DOA information through cyclic position encoding to obtain DOA vector data, and embed the DOA vector data into the mixed voice data through clue encoding to obtain the voice data to be processed.
[0104] As Figure 8 shown, Figure 8 it is a schematic flowchart of a regional voice pickup method provided by an embodiment of the present application. The noisy multi-channel voice data 401 undergoes short-time Fourier transform STFT and time convolutional layer T-ConV to obtain mixed voice data.
[0105] The cyclic position encoding Cyc-pos embeding is used to process the DOA information 403 to generate low-dimensional and continuous DOA vector data. This DOA vector data is broadcast along the time dimension and then passes through a clue encoder composed of a linear layer, layer normalization (LN), and parametric rectified linear unit (PReLU) (see the structure of the clue encoder in Figure 8 the lower right corner). The encoded DOA vector is embedded into the outputs of the time convolutional layer T-Conv and the narrowband layer Narrowband Layer (except the last layer) by element-wise multiplication to enhance the perception of the target speaker's direction and spatial selectivity, and help to focus on the voice signal in the target area in a complex environment.
[0106] S403. Determine the beam width range based on the beam width, filter the voice data in the voice data to be processed except for the voice data corresponding to the beam width range, and obtain the noise-reduced voice data.
[0107] In this embodiment, the beam width range corresponding to the beam width information 404 is encoded, and a range mask is generated through a beam width convolution module BW-convModule including a one-hot encoding, a linear layer, and a 1×1 Conv2D module. The structure of the beam width convolution module BW-convModule is shown in the lower right corner.
[0108] The range mask corresponding to the beam width range can dynamically adjust the desired beam width within the beam width range, filtering out the noise outside the speech data corresponding to the desired beam width in the speech data to be processed. At the same time, the residual connection is used to retain the key information in the speech data to be processed as the original signal, ensuring that the important features of the speech information of the target speaker in the speech data to be processed are not lost while suppressing the noise, and further improving the focusing ability on the speech in the target area.
[0109] S404. Extract the spectral features of the noise-reduced speech data through the crossband layer, and extract the time-frequency features of the noise-reduced speech data through the narrowband layer.
[0110] The noise-reduced speech data passes through the crossband layer Crossband Layer and the narrowband layer Narrowband Layer in sequence. The spectral features of the noise-reduced speech data are extracted through the crossband layer, and the time-frequency features of the noise-reduced speech data are extracted through the narrowband layer.
[0111] As Figure 9 shown, Figure 9 is a schematic structural diagram of a broadband layer and a narrowband layer provided by an embodiment of the present application.
[0112] The crossband layer Crossband Layer is composed of two frequency convolution modules FConv module and a full-band linear module Full-band linear module, and processes each time frame independently. The frequency convolution module FConv module uses grouped convolution FGConv1d along the frequency axis to model local spectral dependencies and capture the correlations between different frequencies. The full-band linear module first reduces the hidden channels from c to C, then processes each hidden channel separately through a group of frequency-direction linear layers, and finally restores the channel dimension to c through a linear layer. Through such a structure, the crossband layer can effectively learn the spectral features of the noise-reduced speech data and improve the analysis and processing ability of different frequency components.
[0113] The Narrowband Layer aims to capture the time dependencies of each frequency and consists of a multi-head self-attention module (MHSA module) and a temporal convolutional feed-forward network (T-ConvFFN module). The multi-head self-attention module (MHSA module) calculates the spatial similarities within each frequency, which helps to separate the speech components from different directions. The temporal convolutional feed-forward network (T-ConvFFN module) enhances the temporal modeling ability by inserting temporal convolutional layers (T-Convs) between two linear transformations, enabling better capture of the temporal dynamic features in the noise-reduced speech data.
[0114] S405. Analyze the noise-reduced speech data based on spectral features and time-frequency features to obtain the speech information of the target speaker in the target area.
[0115] Analyze the noise-reduced speech data based on spectral features and time-frequency features, and map the processed speech information to the estimated target STFT coefficients through a linear output layer (linear). Then, apply the inverse short-time Fourier transform (iSTFT) to convert the speech information, which is a frequency-domain signal, back to the time domain, reconstructing the enhanced and extracted speech information of the target speaker, and completing the extraction task of the speech information of the target speaker in the target area of the noisy multi-channel speech data 401.
[0116] In one embodiment, analyze the noise-reduced speech data based on spectral features and time-frequency features to obtain the initial speech information of the to-be-determined speaker in the to-be-determined area; combine the initial speech information and the mixed speech data, and perform filtering based on DOA vector data embedding and speech data corresponding to the beam width range again to obtain the cyclic noise-reduced speech data; extract the spectral features of the cyclic noise-reduced speech data through the cross-band layer, and extract the time-frequency features of the cyclic noise-reduced speech data through the narrowband layer; analyze the cyclic noise-reduced speech data based on the spectral features and time-frequency features of the cyclic noise-reduced speech data, and repeat the above steps until the end condition is met, then obtain the speech information of the target speaker in the target area.
[0117] In other words, in this embodiment, perform the first round of speech extraction task on the mixed speech data based on DOA information and beam width, so as to obtain the initial speech information of the to-be-determined speaker in the to-be-determined area. Perform the second round of speech extraction task based on the initial speech information, and repeat the extraction multiple times until the speech information of the target speaker in the target area is obtained.
[0118] In this embodiment, the module including the beam broadband convolution module BW-convModule, the crossband layer CrossbandLayer, and the narrowband layer Narrowband Layer is repeated L times. The crossband layer and the narrowband layer are interleaved with each other to jointly enhance the ability of the regional sound pickup deep learning network to distinguish the target area in the mixed speech and extract the speech information of the target speaker, while effectively suppressing noise, reverberation, and interference.
[0119] In this application, for the collected mixed speech data, the DOA information is encoded and embedded into the mixed speech data to obtain the speech data to be processed. By embedding the DOA information, the perception of the direction of the target speaker and the spatial selectivity are enhanced, which helps to focus on the speech information in the target area in a noisy environment and achieve dynamic locking of the target area. Further, according to the specific value of the beam width, the beam width range is dynamically adjusted to filter the speech data in the speech data to be processed except for the speech data corresponding to the beam width range, and the noise-reduced speech data is obtained. The size of the required beam width range can be adjusted according to the actual application scenario, so as to adjust the spatial range of speech extraction, perform precise noise reduction on the speech data to be processed, improve the clarity and accuracy of extracting speech information from the noise-reduced speech data, and provide strong support for subsequent tasks such as speech recognition.
[0120] The following is the device embodiment of this application, which can be used to execute the method embodiment of this application. For the details not disclosed in the device embodiment of this application, please refer to the method embodiment of this application.
[0121] Please refer to Figure 9 , which shows the structural schematic diagram of the regional sound pickup device provided by an exemplary embodiment of this application. The regional sound pickup device can be implemented as all or part of the device through software, hardware, or a combination of both. The device includes an information acquisition module 601, a vector embedding module 602, a data filtering module 603, and a speech extraction module 604.
[0122] The information acquisition module 601 is used to acquire mixed speech data, direction of arrival DOA information, and beam width;
[0123] The vector embedding module 602 is used to perform low-dimensional encoding on the DOA information to obtain DOA vector data, and embed the DOA vector data into the mixed speech data to obtain the speech data to be processed;
[0124] The data filtering module 603 is used to determine the beam width range based on the beam width, and filter the speech data in the speech data to be processed except for the speech data corresponding to the beam width range to obtain the noise-reduced speech data;
[0125] A voice extraction module 604 is configured to perform voice extraction on the noise-reduced voice data to obtain voice information of a target speaker within a target area.
[0126] In one embodiment, the vector embedding module 602 includes:
[0127] A first embedding unit is configured to perform low-dimensional encoding on the DOA information through cyclic position encoding to obtain DOA vector data, and embed the DOA vector data into the mixed voice data to obtain voice data to be processed.
[0128] In one embodiment, the vector embedding module 602 includes:
[0129] A second embedding unit is configured to encode the DOA information to obtain DOA vector data, and embed the DOA vector data into the mixed voice data through clue encoding to obtain voice data to be processed.
[0130] In one embodiment, the data filtering module 603 includes:
[0131] A first filtering unit is configured to determine a beam width range associated with the beam width based on the association relationship between a preset beam width and a beam width range according to the beam width.
[0132] A second filtering unit is configured to encode the beam width range to obtain a range mask, and filter the voice data other than the voice data corresponding to the beam width range in the voice data to be processed according to the range mask to obtain noise-reduced voice data.
[0133] In one embodiment, the second filtering unit includes:
[0134] A first filtering subunit is configured to encode the beam width range to obtain a range mask, and filter the voice data other than the voice data corresponding to the beam width range in the voice data to be processed according to the range mask to obtain voice data to be processed and noise-reduced.
[0135] A second filtering subunit is configured to perform residual connection processing on the voice data to be processed and noise-reduced and the voice data to be processed to obtain noise-reduced voice data.
[0136] In one embodiment, the area sound pickup device further includes:
[0137] A first pre-training module is configured to obtain a pre-trained beam width, pre-trained mixed voice data, and voice information of a speaker in the pre-trained mixed data.
[0138] A second pre-training module, configured to allocate an initial beamwidth range for the pre-trained beamwidth, and extract initial voice information from the pre-trained mixed voice data based on the initial beamwidth range;
[0139] A third pre-training module, configured to adjust the initial beamwidth range allocated for the pre-trained beamwidth according to the initial voice information and the voice information, until the voice information extracted from the pre-trained mixed voice data based on the beamwidth range corresponding to the pre-trained beamwidth meets the training conditions, determine the beamwidth range corresponding to the pre-trained beamwidth, and construct an association relationship between the beamwidth and the beamwidth range.
[0140] In one embodiment, the voice extraction module 604 includes:
[0141] A first extraction unit, configured to extract the spectral features of the noise-reduced voice data through a cross-band layer, and extract the time-frequency features of the noise-reduced voice data through a narrow-band layer;
[0142] A second extraction unit, configured to analyze the noise-reduced voice data based on the spectral features and time-frequency features of the noise-reduced voice data, and obtain the voice information of the target speaker in the target area.
[0143] In one embodiment, the voice extraction module 604 includes:
[0144] A first loop unit, configured to analyze the noise-reduced voice data based on the spectral features and the time-frequency features, and obtain the initial voice information of the to-be-determined speaker in the to-be-determined area;
[0145] A second loop unit, configured to combine the initial voice information and the mixed voice data, and perform filtering based on the DOA vector data embedding and the voice data corresponding to the beamwidth range again to obtain loop noise-reduced voice data;
[0146] A third loop unit, configured to extract the spectral features of the loop noise-reduced voice data through the cross-band layer, and extract the time-frequency features of the loop noise-reduced voice data through the narrow-band layer;
[0147] A fourth loop unit, configured to analyze the loop noise-reduced voice data based on the spectral features and time-frequency features of the loop noise-reduced voice data, and repeat the above steps until the end condition is met, and obtain the voice information of the target speaker in the target area.
[0148] In one embodiment, the area sound pickup device further includes:
[0149] A beam matching module, configured to obtain a beamwidth matching the pickup scenario according to the pickup scenario.
[0150] In this application, for the collected mixed voice data, the DOA information is encoded and embedded into the mixed voice data to obtain the voice data to be processed. By embedding the DOA information, the perception of the direction of the target speaker and the spatial selectivity are enhanced, which helps to focus on the voice information in the target area in a noisy environment and realizes the dynamic locking of the target area. Further, the beam width range is dynamically adjusted according to the specific value of the beam width, and the voice data other than the voice data corresponding to the beam width range in the voice data to be processed is filtered to obtain the noise-reduced voice data. The size of the required beam width range can be adjusted according to the actual application scenario, so as to adjust the spatial range of voice extraction, perform precise noise reduction on the voice data to be processed, improve the clarity and accuracy of extracting voice information from the noise-reduced voice data, and provide strong support for subsequent tasks such as voice recognition.
[0151] It should be noted that when the area sound pickup device provided in the above embodiment executes the area sound pickup method, only the above division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the area sound pickup device provided in the above embodiment and the area sound pickup method embodiment belong to the same concept, and the implementation process is shown in the method embodiment, which will not be elaborated here.
[0152] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0153] The embodiment of the present application also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the area sound pickup method as described in the above Figure 1 - Figure 9 shown embodiments. The specific execution process can refer to the specific description of the Figure 1 - Figure 9 shown embodiments, which will not be elaborated here.
[0154] The present application also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by a processor to perform the area sound pickup method as described in the above Figure 1 - Figure 9 shown embodiments. The specific execution process can refer to the specific description of the Figure 1 - Figure 9 shown embodiments, which will not be elaborated here.
[0155] Please refer to Figure 11 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 11As shown, the electronic device 700 may include: at least one processor 701, at least one network interface 704, a user interface 703, a memory 705, and at least one communication bus 702.
[0156] Among them, the communication bus 702 is used to implement connection communication between these components.
[0157] Among them, the user interface 703 may include a display screen and a camera. Optionally, the user interface 703 may further include a standard wired interface and a wireless interface.
[0158] Among them, the network interface 704 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface).
[0159] Among them, the processor 701 may include one or more processing cores. The processor 701 connects various parts within the entire server 700 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 705, and by calling data stored in the memory 705, the processor 701 performs various functions of the server 700 and processes data. Optionally, the processor 701 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 701 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 701 and may be implemented separately by a single chip.
[0160] Among them, the memory 705 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory 705 includes a non-transitory computer-readable storage medium. The memory 705 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 705 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above method embodiments, etc.; the data storage area may store data involved in the above method embodiments. Optionally, the memory 705 may also be at least one storage device located far from the aforementioned processor 701. As Figure 11 shown, in the memory 705 as a computer storage medium, an operating system, a network communication module, a user interface module, and a regional voice pickup application program may be included.
[0161] In Figure 11 in the electronic device 700 shown, the user interface 703 is mainly used to provide an input interface for the user to obtain user input data; while the processor 701 can be used to call the regional voice pickup application program stored in the memory 705 and specifically perform the following operations:
[0162] Obtain mixed voice data, direction of arrival (DOA) information, and beam width;
[0163] Encode the DOA information to obtain DOA vector data, and embed the DOA vector data into the mixed voice data to obtain voice data to be processed;
[0164] Determine a beam width range based on the beam width, and filter the voice data in the voice data to be processed except for the voice data corresponding to the beam width range to obtain noise-reduced voice data;
[0165] Perform voice extraction on the noise-reduced voice data to obtain the voice information of the target speaker in the target area.
[0166] In one embodiment, when the processor 701 executes encoding the DOA information to obtain DOA vector data and embedding the DOA vector data into the mixed voice data to obtain voice data to be processed, it specifically executes:
[0167] Low - dimensionally encode the DOA information through cyclic position encoding to obtain DOA vector data, and embed the DOA vector data into the mixed speech data to obtain the speech data to be processed.
[0168] In one embodiment, the processor 701 executes encoding the DOA information to obtain DOA vector data, and embedding the DOA vector data into the mixed speech data to obtain the speech data to be processed, and specifically executes:
[0169] Encode the DOA information to obtain DOA vector data, and embed the DOA vector data into the mixed speech data through clue encoding to obtain the speech data to be processed.
[0170] In one embodiment, the processor 701 executes determining a beam width range based on the beam width, and filtering speech data in the speech data to be processed other than the speech data corresponding to the beam width to obtain noise - reduced speech data, and specifically executes:
[0171] Based on the association relationship between the preset beam width and the beam width range, determine the beam width range associated with the beam width based on the beam width;
[0172] Encode the beam width range to obtain a range mask, and filter speech data in the speech data to be processed other than the speech data corresponding to the beam width range according to the range mask to obtain noise - reduced speech data.
[0173] In one embodiment, the processor 701 executes encoding the beam width range to obtain a range mask, and filtering speech data in the speech data to be processed other than the speech data corresponding to the beam width range according to the range mask to obtain noise - reduced speech data, and specifically executes:
[0174] Encode the beam width range to obtain a range mask, and filter speech data in the speech data to be processed other than the speech data corresponding to the beam width range according to the range mask to obtain the to - be - processed noise - reduced speech data;
[0175] Perform residual connection processing on the to - be - processed noise - reduced speech data and the speech data to be processed to obtain noise - reduced speech data.
[0176] In one embodiment, before the processor 701 executes obtaining the mixed speech data, the direction of arrival DOA information, and the beam width, it also executes:
[0177] Obtain the pre - trained beam width and pre - trained mixed speech data, and the speech information of the speaker in the pre - trained mixed data;
[0178] Allocate an initial beamwidth range for the pre-training beamwidth, and extract initial speech information from the pre-training mixed speech data based on the initial beamwidth range;
[0179] According to the initial speech information and the speech information, adjust the initial beamwidth range allocated for the pre-training beamwidth until the speech information extracted from the pre-training mixed speech data based on the beamwidth range corresponding to the pre-training beamwidth meets the training conditions, then determine the beamwidth range corresponding to the pre-training beamwidth, and construct the association relationship between the beamwidth and the beamwidth range.
[0180] In one embodiment, the processor 701 executes the speech extraction on the noise-reduced speech data to obtain the speech information of the target speaker in the target area, and specifically executes:
[0181] Extract the spectral features of the noise-reduced speech data through the cross-band layer, and extract the time-frequency features of the noise-reduced speech data through the narrow-band layer;
[0182] Analyze the noise-reduced speech data based on the spectral features and time-frequency features of the noise-reduced speech data to obtain the speech information of the target speaker in the target area.
[0183] In one embodiment, the processor 701 executes the speech extraction on the noise-reduced speech data to obtain the speech information of the target speaker in the target area, and specifically executes:
[0184] Analyze the noise-reduced speech data based on the spectral features and the time-frequency features to obtain the initial speech information of the to-be-determined speaker in the to-be-determined area;
[0185] Combine the initial speech information and the mixed speech data, and perform filtering based on the DOA vector data embedding and the speech data corresponding to the beamwidth range again to obtain the cyclic noise-reduced speech data;
[0186] Extract the spectral features of the cyclic noise-reduced speech data through the cross-band layer, and extract the time-frequency features of the cyclic noise-reduced speech data through the narrow-band layer;
[0187] Analyze the cyclic noise-reduced speech data based on the spectral features and time-frequency features of the cyclic noise-reduced speech data, and repeat the above steps until the end condition is met, then obtain the speech information of the target speaker in the target area.
[0188] In one embodiment, before the processor 701 executes the obtaining of the mixed speech data, the direction of arrival DOA information, and the beamwidth, it also executes:
[0189] According to the pick-up scenario, obtain the beamwidth matching the pick-up scenario.
[0190] In the present application, for the collected mixed voice data, the DOA information is encoded and embedded into the mixed voice data to obtain the voice data to be processed. By embedding the DOA information, the perception of the direction of the target speaker and the spatial selectivity are enhanced, which helps to focus on the voice information in the target area in a noisy environment and realizes the dynamic locking of the target area. Further, the beam width range is dynamically adjusted according to the specific value of the beam width, and the voice data other than the voice data corresponding to the beam width range in the voice data to be processed is filtered to obtain the noise-reduced voice data. The size of the required beam width range can be adjusted according to the actual application scenario, so as to adjust the spatial range of voice extraction, perform precise noise reduction on the voice data to be processed, improve the clarity and accuracy of extracting voice information from the noise-reduced voice data, and provide strong support for subsequent tasks such as voice recognition.
[0191] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory or a random access memory, etc.
[0192] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0193] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A regional sound pickup method, characterized in that: The method comprises: Obtain mixed voice data, direction of arrival DOA information and beam width; Encoding the DOA information to obtain DOA vector data, and embedding the DOA vector data into the mixed voice data to obtain voice data to be processed; Determine a beam width range based on the beam width, filter the voice data to be processed except the voice data corresponding to the beam width range, and obtain noise-reduced voice data; Speech extraction is performed on the noise-reduced speech data to obtain speech information of a target speaker in a target area.
2. The regional sound pickup method according to claim 1, characterized in that: The step of encoding the DOA information to obtain DOA vector data, and embedding the DOA vector data into the mixed voice data to obtain voice data to be processed, comprises: The DOA information is low-dimensionally encoded through cyclic position coding to obtain DOA vector data, and the DOA vector data is embedded into the mixed speech data to obtain the speech data to be processed.
3. The regional sound pickup method according to claim 1 or 2, characterized in that: The step of encoding the DOA information to obtain DOA vector data, and embedding the DOA vector data into the mixed voice data to obtain voice data to be processed, comprises: The DOA information is encoded to obtain DOA vector data, and the DOA vector data is embedded into the mixed speech data through clue encoding to obtain speech data to be processed.
4. The regional sound pickup method according to claim 1, characterized in that: The step of determining a beam width range based on the beam width, filtering the speech data to be processed except the speech data corresponding to the beam width, and obtaining the noise-reduced speech data comprises: According to a preset association relationship between the beam width and the beam width range, determining a beam width range associated with the beam width based on the beam width; The beam width range is encoded to obtain a range mask, and voice data other than voice data corresponding to the beam width range in the voice data to be processed is filtered according to the range mask to obtain noise-reduced voice data.
5. The regional sound pickup method according to claim 4, characterized in that: The encoding of the beam width range to obtain a range mask, and filtering the voice data to be processed except the voice data corresponding to the beam width range according to the range mask to obtain the noise-reduced voice data, includes: Encoding the beam width range to obtain a range mask, and filtering the voice data to be processed except the voice data corresponding to the beam width range according to the range mask to obtain the noise-reduced voice data to be processed; Residual connection processing is performed on the noise reduction speech data to be processed and the speech data to be processed to obtain noise reduction speech data.
6. The regional sound pickup method according to claim 1 or 4, characterized in that: Before obtaining the mixed voice data, the direction of arrival DOA information and the beam width, the method further includes: Obtaining pre-trained beam width and pre-trained mixed speech data, as well as speech information of the speaker in the pre-trained mixed data; Allocating an initial beam width range to the pre-trained beam width, and extracting initial speech information from the pre-trained mixed speech data based on the initial beam width range; According to the initial voice information and the voice information, the initial beam width range allocated to the pre-training beam width is adjusted until the voice information extracted from the pre-training mixed voice data based on the beam width range corresponding to the pre-training beam width meets the training conditions, the beam width range corresponding to the pre-training beam width is determined, and the correlation relationship between the beam width and the beam width range is constructed.
7. The regional sound pickup method according to claim 1, characterized in that: The performing voice extraction on the noise reduction voice data to obtain voice information of a target speaker in a target area includes: Extracting frequency spectrum features of the noise-reduced speech data through a cross-band layer, and extracting time-frequency features of the noise-reduced speech data through a narrow-band layer; The noise reduction speech data is analyzed based on the frequency spectrum characteristics and time-frequency characteristics of the noise reduction speech data to obtain speech information of a target speaker in a target area.
8. The regional sound pickup method according to claim 7, characterized in that: The step of analyzing the noise reduction speech data based on the spectral characteristics and time-frequency characteristics of the noise reduction speech data to obtain speech information of a target speaker in a target area includes: Analyze the noise reduction speech data based on the frequency spectrum feature and the time-frequency feature pair to obtain initial speech information of the speaker to be determined in the to-be-determined area; Combining the initial voice information and the mixed voice data, again performing filtering based on the DOA vector data embedding and based on the voice data corresponding to the beam width range to obtain cyclic noise reduction voice data; Extracting frequency spectrum features of the cyclic denoised speech data through the cross-band layer, and extracting time-frequency features of the cyclic denoised speech data through the narrowband layer; The cyclic denoised speech data is analyzed based on the spectrum characteristics and time-frequency characteristics of the cyclic denoised speech data, and the above steps are repeated until an end condition is met to obtain the speech information of the target speaker in the target area.
9. The regional sound pickup method according to claim 1, characterized in that: Before obtaining the mixed voice data, the direction of arrival DOA information and the beam width, the method further includes: According to the sound pickup scene, a beam width matching the sound pickup scene is acquired.
10. A regional sound pickup device, characterized in that: The device comprises: An information acquisition module is used to obtain mixed voice data, direction of arrival DOA information and beam width; A vector embedding module, used for performing low-dimensional encoding on the DOA information to obtain DOA vector data, and embedding the DOA vector data into the mixed voice data to obtain voice data to be processed; A data filtering module, configured to determine a beam width range based on the beam width, filter the voice data to be processed except the voice data corresponding to the beam width range, and obtain noise-reduced voice data; The speech extraction module is used to perform speech extraction on the noise reduction speech data to obtain speech information of a target speaker in a target area.
11. A computer storage medium, characterized in that: The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 9.
12. A computer program product, characterized in that The computer program product stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 9.
13. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 9.