Method, device and equipment for training sound source separation model
By acquiring sound signals in diverse scenarios, determining room impulse response functions, and fine-tuning a deep neural network model, the method addresses the challenges of speech separation in smart glasses, achieving improved speech signal separation and user experience.
Patent Information
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- SHANGHAI QIANWEN ZHILIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-17
AI Technical Summary
Smart glasses face challenges in accurately separating wearer's speech from ambient speech due to complex and ever-changing acoustic environments, with existing methods suffering from low simulation accuracy, poor spatial resolution, slow convergence, and limited adaptability, leading to poor speech signal separation results.
A method involving the acquisition of sound signals in different scenarios and paths, determination of room impulse response functions, generation of basic training data, and fine-tuning a deep neural network model using real-world scene data to generate a target separation model.
The method effectively improves speech signal separation in smart glasses, enhancing user experience by providing a high-quality acoustic dataset and adaptable speech separation performance across various environments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202511895965.6 (22) Application Date 2025.12.15 (71) Applicant Shanghai Qianwen Zhilian Artificial Intelligence Technology Co., Ltd. Address Room 407-4, 4th Floor, Building X7, No.8 Longyao Road, Xuhui District, Shanghai 200232 (72) Inventors Xue Zheng, Liu Chengwei, Liang Xiaotao, Xue Shaofei (74) Patent Agency Beijing Ruipai Intellectual Property Agency Co., Ltd. 11597 Patent Attorney Liu Feng, Yang Chunxiao (51) Int.Cl. G10L 21 / 028 (2013.01) G10L 21 / 0208 (2013.01) G10L 25 / 30 (2013.01) (54) Invention Title: A Method, Apparatus, and Device for Training a Sound Source Separation Model (57) Abstract: This invention discloses a method, apparatus, and device for training a sound source separation model. In this embodiment, multiple sound signals from a wearer and / or a non-wearer are acquired. These multiple sound signals are generated by a smart device recording playback signals in different scenarios and paths. The playback signals are sound signals emitted by the wearer and / or the non-wearer. Multiple room impulse response (RIR) functions are determined based on the multiple sound signals. Multiple basic training data are generated based on the multiple RIR functions. A basic separation model is trained based on the multiple basic training data. Multiple real-world scene data from the wearer and non-wearer are acquired. The basic separation model is fine-tuned based on the multiple real-world scene data to generate a target separation model. Through the above method, a target separation model is generated, effectively improving the speech signal separation effect and enhancing the user experience. Claims 2 pages, Description 11 pages, Drawings 8 pages, CN 121583283 A 2026.02.27 CN 1 21 58 32 83 A 1. A method for training a sound source separation model, characterized in that the method includes: acquiring multiple sound signals from a wearer and / or a non-wearer, wherein the multiple sound signals are generated by a smart device recording playback signals under different scenarios and paths, and the playback signals are sound signals emitted by the wearer and / or the non-wearer; determining multiple room impulse response (RIR) functions based on the multiple sound signals; generating multiple basic training data based on the multiple RIR functions; training a basic separation model based on the multiple basic training data; acquiring multiple real-world scenario data of the wearer and non-wearer; fine-tuning the basic separation model based on the multiple real-world scenario data to generate a target separation model. 2. The method according to claim 1, characterized in that determining multiple room impulse response (RIR) functions based on the multiple sound signals specifically includes:The method aligns the plurality of sound signals with the playback signal and performs time correction on the plurality of sound signals; it then performs frequency sweeping and signal quality inspection on the time-corrected plurality of sound signals to determine a plurality of normal sound signals; and finally performs deconvolution calculation on the plurality of normal generated sound signals to generate a plurality of RIR functions. 3. The method according to claim 1, wherein the method further includes: naming and saving the RIR functions according to a set data naming format. 4. The method according to claim 1, wherein generating a plurality of basic training data based on the plurality of RIR functions specifically includes: processing single-channel clean speech or multi-channel speech respectively through the plurality of RIR functions to generate basic training data, wherein the basic training data includes the input data and output data of the basic separation model. 5. The method according to claim 4, wherein the method further includes: superimposing real noise to generate the basic training data. 6. The method according to claim 1, wherein the basic separation model is a deep neural network model. 7. The method according to claim 1, wherein acquiring multiple real-world scene data of the wearer and non-wearer specifically includes: acquiring multiple quiet scene voice signals of the wearer and non-wearer; superimposing the multiple quiet scene voice signals with multiple noisy scene signals according to a set signal-to-noise ratio and signal-to-interference ratio to generate the multiple real-world scene data. 8. The method according to claim 1, wherein the method further includes: acquiring at least two RIR functions from the multiple room impulse response RIR functions that match the business requirements, based on the scene range of the wearer and / or non-wearer in the business requirements. 9. The method according to claim 1, wherein the different scenes include anechoic chamber scenes, conference room scenes, office area scenes, and outdoor scenes; the different paths include the voice path of the wearer and the voice path of the non-wearer. 10. An apparatus for training a sound source separation model, characterized in that the apparatus comprises: an acquisition unit, configured to acquire multiple sound signals from a wearer and / or a non-wearer, wherein the multiple sound signals are generated by a smart device recording playback signals under different scenarios and paths, and the playback signals are sound signals emitted by the wearer and / or the non-wearer; a determination unit, configured to determine multiple room impulse response (RIR) functions based on the multiple sound signals; a generation unit, configured to generate multiple basic training data based on the multiple RIR functions; and a training unit, configured to train a basic separation model based on the multiple basic training data; the acquisition unit is further configured to: acquire multiple real-world scene data of the wearer and non-wearer;A fine-tuning unit is used to fine-tune the basic separation model based on the multiple real-world scene data to generate a target separation model. 11. An electronic device, comprising a memory and a processor, characterized in that the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-9. 12. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method as described in any one of claims 1-9. Claims 2 / 2 Page 3 CN 121583283 A Method, Apparatus and Device for Training a Sound Source Separation Model Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to a method, apparatus and device for training a sound source separation model. Background Technology
[0002] With the rapid development of Artificial Intelligence (AI) and smart wearable device technology, smart glasses, as a new generation of human-computer interaction carrier, are gradually moving from concept to practical application. Smart glasses integrate a variety of cutting-edge technologies such as speech recognition, computer vision, and augmented reality, aiming to provide users with a more natural and convenient interactive experience. Among them, the voice interaction function, as the most intuitive human-computer interface, has become an important part of the core technology of smart glasses. In voice interaction technology, the accuracy of speech signal separation and recognition directly determines the user experience. Smart glasses face the challenge of complex and ever-changing acoustic environment in actual use. Unlike traditional fixed voice devices, smart glasses need to process multiple sound source signals from different directions and distances while in motion, including the wearer's own voice commands, conversations in the surrounding environment, background noise, etc. The above-mentioned complex and ever-changing acoustic environment makes it difficult for traditional voice processing methods to meet the application needs of smart glasses.
[0003] In the prior art, speech signal separation is performed using speech processing systems based on simple room impulse response (RIR) simulation, traditional beamforming speech enhancement systems, and traditional speech processing methods based on blind source separation. However, the speech processing systems based on simple RIR simulation suffer from drawbacks such as low simulation accuracy, insufficient realism, and weak generalization ability. The traditional beamforming speech enhancement systems suffer from drawbacks such as low spatial resolution, poor dynamic adaptability, weak ability to handle complex environments, and limited algorithm capabilities. The traditional speech processing methods based on blind source separation suffer from drawbacks such as slow convergence speed, uncertainty in the arrangement of the separated signals, high computational complexity, and limited algorithm capabilities. These drawbacks lead to poor speech signal separation results.
[0004] In summary, how to improve the speech signal separation effect is a problem that needs to be solved.
[0005] In view of the above, embodiments of the present invention provide a method, apparatus, and device for training a sound source separation model, generating a target separation model, effectively improving the speech signal separation effect, and enhancing the user experience.
[0006] In a first aspect, embodiments of the present invention provide a method for training a sound source separation model, the method comprising: acquiring multiple sound signals from a wearer and / or a non-wearer, wherein the multiple sound signals are generated by a smart device recording playback signals in different scenarios and different paths, and the playback signals are sound signals emitted by the wearer and / or the non-wearer; determining multiple room impulse response (RIR) functions based on the multiple sound signals; generating multiple basic training data based on the multiple RIR functions; training a basic separation model based on the multiple basic training data; acquiring multiple real-world scene data of the wearer and non-wearer; and fine-tuning the basic separation model based on the multiple real-world scene data to generate a target separation model.
[0007] Optionally, determining multiple room impulse response (RIR) functions based on the multiple sound signals specifically includes: aligning the multiple sound signals with the playback signal, and performing time correction on the multiple sound signals; performing frequency sweeping and signal quality inspection on the time-corrected multiple sound signals to determine multiple normal sound signals; and performing deconvolution calculation on the multiple normal generated sound signals to generate multiple RIR functions.
[0008] Optionally, the method further includes: naming and saving the RIR functions according to a set data naming format.
[0009] Optionally, generating multiple basic training data based on the multiple RIR functions specifically includes: processing single-channel clean speech or multi-channel speech through the multiple RIR functions to generate basic training data, wherein the basic training data includes the input data and output data of the basic separation model.
[0010] Optionally, the method further includes: superimposing real noise to generate the basic training data.
[0011] Optionally, the basic separation model is a deep neural network model.
[0012] Optionally, acquiring multiple real-world scene data of the wearer and non-wearer specifically includes: acquiring multiple quiet scene speech signals of the wearer and non-wearer; superimposing the multiple quiet scene speech signals with multiple noisy scene signals according to a set signal-to-noise ratio and signal-to-interference ratio to generate the multiple real-world scene data.
[0013] Optionally, the method further includes: acquiring at least two RIR functions from the multiple room impulse response RIR functions that match the business requirements based on the scene range of the wearer and / or non-wearer in the business requirements.
[0014] Optionally, the different scenes include anechoic chamber scenes, conference room scenes, office area scenes, and outdoor scenes;The different paths include the voice path of the wearer and the voice path of the non-wearer.
[0015] In a second aspect, embodiments of the present invention provide an apparatus for training a sound source separation model, the apparatus comprising: an acquisition unit, configured to acquire multiple sound signals of a wearer and / or a non-wearer, wherein the multiple sound signals are generated by a smart device recording playback signals under different scenarios and different paths, and the playback signals are sound signals emitted by the wearer and / or the non-wearer; a determination unit, configured to determine multiple room impulse response (RIR) functions based on the multiple sound signals; a generation unit, configured to generate multiple basic training data based on the multiple RIR functions; a training unit, configured to train a basic separation model based on the multiple basic training data; the acquisition unit is further configured to: acquire multiple real-world scene data of the wearer and the non-wearer; and a fine-tuning unit, configured to fine-tune the basic separation model based on the multiple real-world scene data to generate a target separation model.
[0016] Optionally, the determining unit is specifically used for: aligning the plurality of sound signals with the playback signal, and performing time correction on the plurality of sound signals; performing frequency sweeping and signal quality inspection on the time-corrected plurality of sound signals to determine a plurality of normal sound signals; and performing deconvolution calculation on the plurality of normal generated sound signals to generate a plurality of RIR functions.
[0017] Optionally, the device further includes: a saving unit, used to name and save the RIR functions according to a set data naming format.
[0018] Optionally, the generating unit is further used for: processing single-channel clean speech or multi-channel speech through the plurality of RIR functions to generate basic training data, wherein the basic training data includes the input data and output data of the basic separation model.
[0019] Optionally, the generating unit is further used for: superimposing real noise to generate the basic training data.
[0020] Optionally, the basic separation model is a deep neural network model.
[0021] Optionally, the acquisition unit is further configured to: acquire multiple quiet scene voice signals of the wearer and the non-wearer; superimpose the multiple quiet scene voice signals with multiple noisy scene signals according to a set signal-to-noise ratio and signal-to-interference ratio to generate the multiple real scene data.
[0022] Optionally, the acquisition unit is further configured to: acquire at least two RIR functions from the multiple room impulse response RIR functions that match the business requirements according to the scene range of the wearer and / or non-wearer in the business requirements. Specification 2 / 11 Page 5 CN 121583283 A
[0023] Optionally, the different scenes include anechoic chamber scene, conference room scene, office area scene and outdoor scene; the different paths include the voice path of the wearer and the voice path of the non-wearer.
[0024] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect or any one of the first aspects.
[0025] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the method as described in the first aspect or any one of the first aspects.
[0026] In embodiments of the present invention, multiple sound signals from a wearer and / or a non-wearer are acquired, wherein the multiple sound signals are generated by a smart device recording playback signals in different scenarios and different paths, and the playback signals are sound signals emitted by the wearer and / or the non-wearer; multiple room impulse response (RIR) functions are determined based on the multiple sound signals; multiple basic training data are generated based on the multiple RIR functions; a basic separation model is trained based on the multiple basic training data; multiple real-world scene data of the wearer and non-wearer are acquired; and a target separation model is generated by fine-tuning the basic separation model based on the multiple real-world scene data. By using the above method, a target separation model is generated using multi-dimensional training data, which effectively improves the speech signal separation effect and enhances the user experience. Brief Description of the Drawings
[0027] The above and other objects, features, and advantages of the present invention will become clearer from the following description of embodiments of the present invention with reference to the accompanying drawings, in which: Figure 1 is a flowchart of a method for training a sound source separation model according to an embodiment of the present invention; Figure 2 is a schematic diagram of a smart wearable device and speaker placement according to an embodiment of the present invention; Figure 3 is a flowchart of a method for determining the RIR function according to an embodiment of the present invention; Figure 4 is a flowchart of a method for generating basic training data according to an embodiment of the present invention; Figure 5 is a flowchart of another method for generating basic training data according to an embodiment of the present invention; Figure 6 is a flowchart of another method for training a sound source separation model according to an embodiment of the present invention; Figure 7 is a schematic diagram of a device for training a sound source separation model according to an embodiment of the present invention; Figure 8 is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Description of Embodiments
[0028] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without these detailed descriptions. To avoid obscuring the substance of this application, well-known methods, processes, flows, components, and circuits are not described in detail.
[0029] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0030] Unless the context explicitly requires it, the words "including," "comprising," and similar terms throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to."
[0031] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more.
[0032] In the prior art, smart glasses face several technical problems in speech processing. Problem 1: It is difficult to effectively separate the wearer's speech from the ambient speech. Specifically, existing smart glasses usually use simple energy thresholding or directional filtering methods for speech separation. These methods work well in quiet environments but poorly in complex real-world environments. The main reasons for this problem are the lack of accurate modeling of the wearer's specific acoustic propagation path, the inability to effectively distinguish the acoustic features of the wearer's speech from near-field environmental dialogue, and the neglect of the influence of the glasses wearing state on acoustic signal propagation. Problem 2: Insufficient quality and diversity of training data. Specifically, current mainstream speech separation algorithms rely on a large amount of high-quality training data, but existing datasets have the following limitations: First, most datasets are based on fixed microphone arrays and cannot reflect the acoustic characteristics of smart glasses under wearing conditions; second, there is a lack of specialized data for near-field speech interaction scenarios; third, the data annotation accuracy is limited, making it difficult to support the needs of high-precision separation algorithms; finally... The scenario coverage is incomplete, failing to encompass the various usage environments that smart glasses may encounter; thirdly, the acoustic environment adaptability is poor. Specifically, smart glasses will face drastically different acoustic environments in different usage scenarios; for example, in indoor environments, there are abundant early reflections and reverberation; in outdoor environments, they mainly face wind noise and traffic noise interference; in meeting scenarios, they need to handle the complex situation of multiple people speaking at the same time; in mobile scenarios, there are Doppler effects and wind noise introduced by device movement.
[0033] Currently, speech signal separation is generally performed using the following three methods: Method 1, a speech processing system based on simple room impulse response (RIR) simulation; Method 2, a traditional beamforming speech enhancement system; and Method 3, a traditional speech processing method based on blind source separation. Specifically, Method 1 uses geometric acoustics to simulate room impulse response, trains and tests algorithms based on simulated RIR data, uses a simplified acoustic environment model, and employs basic signal processing algorithms. However, Method 1 suffers from simulation accuracy defects, mainly due to the simplification of physical modeling; the geometric acoustics method ignores complex physical phenomena such as sound wave diffraction and scattering, and the environment is complex.The simulation suffers from several shortcomings: insufficient realism, an overly idealized environment that fails to reflect the complexity of the real world, and a lack of dynamic characteristics, failing to model dynamic environmental changes. Therefore, the application of method one is relatively poor, specifically due to insufficient realism and the resulting RIR values being inadequate. The data differs significantly from real measurement data; the algorithm has poor robustness, and the performance of the algorithm trained on simulation data degrades significantly in real environments; generalization ability is limited, and the limitations of simulation methods restrict the algorithm's scene adaptability; Method 2 specifically uses a microphone array for spatial filtering, and a traditional beamforming algorithm for speech enhancement, using a fixed directional pattern and a simple noise suppression method; the beamforming utilizes the spatial distribution characteristics of the microphone array to form a directional receiving beam through digital signal processing; Method 2 suffers from spatial resolution defects and complex environment processing defects. Specifically, the spatial resolution defects include directional limitations, with traditional beamforming methods having limited spatial resolution, making it difficult to accurately distinguish between similar sound sources; sidelobe interference, with significant sidelobe interference affecting the separation effect; poor dynamic adaptability, unable to effectively track the positional changes of moving sound sources; and complex environment processing defects including insufficient reverberation suppression, resulting in significant performance degradation in reverberant environments; difficulty in processing multiple sound sources, making it difficult to effectively process multiple simultaneously emitting sound sources; and poor noise robustness, performing poorly in non-Gaussian noise environments; Method 3 specifically employs Independent Component Analysis (ICA). Blind source separation algorithms such as ICA (Independent Characteristic Algorithm) are based on the statistical independence assumption for signal separation, using traditional feature extraction methods and fixed separation criteria. However, these methods suffer from theoretical limitations and practical shortcomings. Specifically, the theoretical limitations include: the independence assumption not holding true (in real-world dialogue scenarios, multi-person speech often exhibits correlation, violating the statistical independence assumption); the linear mixing assumption (while mixing processes in real acoustic environments are often non-linear); and the instantaneous mixing assumption (ignoring the time delay effect of sound wave propagation). Practical shortcomings include convergence issues (the ICA algorithm sometimes converges slowly or not at all); and permutation ambiguity (the separated signal may appear unclear). (See page 4 / 11 of the document, CN 121583283 A). There are issues with arrangement uncertainty; computational complexity is high, and traditional blind source separation algorithms have high computational complexity, making it difficult to meet real-time processing requirements; in addition, deep learning solutions have been proposed in existing technologies; in summary, traditional beamforming and blind source separation-based solutions cannot meet the business needs of smart glasses products due to their algorithmic limitations; while deep learning solutions are limited by data quality and quantity, and cannot adequately meet the separation requirements of wearers or non-wearers in various scenarios.
[0034] Therefore, how to improve the speech signal separation effect is a problem that needs to be solved.
[0035] In this embodiment of the invention, in order to solve the above problems, a method for training a sound source separation model is proposed, as shown in Figure 1. The method includes: Step S101, acquiring multiple sound signals from the wearer and / or non-wearer.
[0036] Specifically, the multiple sound signals are generated by the smart device recording playback signals in different scenarios and paths, and the playback signals are sound signals emitted by the wearer and / or non-wearer.
[0037] In one possible implementation, the different scenarios include an anechoic chamber scenario, a conference room scenario, an office area scenario, and an outdoor scenario; wherein, the anechoic chamber scenario is a pure and reflection-free ideal acoustic environment; the conference room scenario represents a typical indoor office environment with a medium reverberation time; the office area scenario simulates an open office environment, containing complex background noise and multiple sound sources; the outdoor scenario represents an outdoor usage environment, mainly facing wind noise and traffic noise interference.
[0038] In one possible implementation, the different paths include the wearer's voice path and the non-wearer's voice path; wherein, the voice path refers to the complete path from the sound source to the sound source, the wearer's voice path uses a smart wearable device to play sound, and uses the same smart wearable device to pick up sound; assuming, the smart wearable device is smart glasses, the wearing posture includes a standard wearing posture and a typical offset posture, wherein, the typical offset posture includes turning the head to the left, turning the head to the right, and wearing it in combination with myopia glasses, etc., different head shapes are selected when wearing smart glasses, and different scenarios are selected under each path to traverse the wearer's voice path; the non-wearer's voice path uses a speaker to play sound, and uses a smart wearable device to pick up sound; assuming, the smart wearable device is smart glasses, the smart glasses are placed in the center, and the speaker is placed at a distance of 0.5 meters (m) from the smart glasses with a 4m dialogue range, covering the near field, In the mid-field and far-field, the speaker is placed in a 360-degree omnidirectional range, as shown in Figure 2. The speaker is placed in each dialogue range and angle range to form a new voice path with the smart wearable device placed in the center. Different scenarios are selected under each path to traverse the voice path of the non-wearer.
[0039] In this embodiment of the invention, by recording voices of the wearer and non-wearer in different scenarios and different paths, the sound signals can be differentiated, which helps the smart wearable device to perform acoustic characteristic analysis in a variety of practical use scenarios. The smart wearable device can be smart glasses, smart headphones, etc.
[0040] Step S102: Determine multiple room impulse response RIR functions based on the multiple sound signals.
[0041] In one possible implementation, the determination of multiple room impulse response RIR functions based on the multiple sound signals is shown in Figure 3, including the following:Step S301: Align the plurality of sound signals with the playback signal and perform time correction on the plurality of sound signals.
[0042] Specifically, the signal alignment is the precise time synchronization of the plurality of sound signals with the playback signal.
[0043] In one possible implementation, the process of acquiring the plurality of sound signals before the signal alignment can also be called frequency sweep recording.
[0044] Step S302: Perform frequency sweeping and interception on the time-corrected plurality of sound signals to acquire sound signals in the effective frequency domain response range. Specification 5 / 11 Page 8 CN 121583283 A
[0045] In one possible implementation, the frequency sweeping and interception can also be called signal extraction.
[0046] Step S303: Perform signal quality inspection on the sound signals in the effective frequency domain response range to determine the plurality of normal sound signals.
[0047] Specifically, the signal quality inspection is used to remove invalid data caused by abnormal operation during the recording process.
[0048] Step S304: Perform deconvolution calculations on the multiple normal generated sound signals to generate multiple RIR functions.
[0049] In this embodiment of the invention, by recording sound signals differently in different scenarios, and through systematic environment selection, path traversal, and signal quality control, a high-quality and diverse acoustic data foundation is provided for subsequent speech separation model training, effectively solving the technical problems of simple RIR data acquisition methods and incomplete scenario coverage.
[0050] Step S103: Generate multiple basic training data based on the multiple RIR functions.
[0051] Specifically, single-channel clean speech or multi-channel speech is processed through the multiple RIR functions to generate basic training data, wherein the basic training data includes the input data and output data of the basic separation model.
[0052] In the process of generating basic training data, real noise can also be superimposed to generate the basic training data.
[0053] The following describes in detail the single-speaker scenario simulation and the multi-speaker scenario simulation through two specific embodiments, as follows: Specific Embodiment 1: Single-speaker scenario simulation, as shown in Figure 4, single-channel clean speech is processed through the multiple RIR functions to generate basic training data; firstly, the clean speech, noise data, and RIR functions are determined, wherein the clean speech, the noise data, and the RIR functions are all forced to be single-channel during loading; then, gain control is determined, and the gain control includes signal-to-noise ratio and volume. Before performing gain control, the clean speech is processed by the RIR function, and the clean speech processed by the RIR function undergoes early RIR and gain control to generate clean speech.The direct speech (speech-drb) is determined as the target speech based on the direct speech of the clean speech. The target speech is the output data of the basic separation model. After processing the clean speech by the RIR function and controlling the gain, the speech reverberation of the clean speech (speech-rvb) is generated. The speech reverberation of the clean speech and mono noise are superimposed to generate mixed speech. The mixed speech is the input data of the basic separation model. The mono noise can also be controlled by gain. After the above processing, a large amount of basic training data can be generated, covering various acoustic conditions.
[0054] In one possible implementation, label information can also be set in each piece of basic training data, such as speaker location, scene, etc.
[0055] Specific embodiment two, multi-speaker scene simulation, as shown in Figure 5, multi-channel speech is processed by the multiple RIR functions to generate basic training data.
[0056] First, the clean speech source noise, RIR function, recorded noise, and interference speech are determined. The clean speech, source noise, and interference speech are all forced to be single-channel during loading. The source noise is single-channel noise. The RIR function is multi-channel during loading. The recorded noise is multi-channel and matched with the RIR function during loading. Then, gain control is determined. The gain control includes signal-to-interference ratio (SIR), signal-to-noise ratio (SNR), and volume. Before gain control, the clean speech is processed by the RIR function. The clean speech processed by the RIR function undergoes early RIR and gain control to generate the clean speech direct tone (speech-drb). Based on the clean speech processed by the RIR function and gain control, clean speech reverberation (speech-rvb) is generated. The recorded noise is bypassed, and the mono noise is upmixed to generate multi-channel noise. The interference speech is processed by the RIR function. (See page 6 / 11 of the RIR function specification, CN 121583283 A) The process involves performing early RIR on the interference speech processed by the RIR function and then applying gain control to generate the interference speech's direct speech tone (disturb-DRB). Based on the clean speech processed by the RIR function and after gain control, a interference speech reverberation tone (disturb-RVB) is generated. The reverberation tone of the clean speech, the multichannel noise, and the interference speech's reverberation tone are then superimposed to generate a mixture. This mixture is the input data of the basic separation model. The multichannel noise can also be controlled by gain. The process is based on the direct speech tone of the clean speech and the reverberation tone of the interference speech.The direct sound is determined as a signal combine, and the combined signal is determined as the target speech, which is the output data of the basic separation model.
[0057] In one possible implementation, label information can also be set in each piece of basic training data, such as speaker location, scene, etc.
[0058] In this embodiment of the invention, a large amount of basic training data can be generated by the above method to determine the balance and representativeness of the basic training data in data distribution; during simulation training, configuration can be performed, and the configuration includes regular training configuration and data path configuration. The regular training configuration includes batch size, number of working processes, etc.; the data path configuration includes speech list file, noise list file and room impulse response list file, etc.
[0059] Step S104: Train the basic separation model according to the multiple basic training data.
[0060] Specifically, the basic separation model is a deep neural network model. The deep neural network architecture of the deep neural network model includes Conv-TasNet, Dual-Path Recurrent Neural Network (DPRNN), etc. The optimization objective function of the basic separation model includes signal distortion ratio and source interference ratio, etc. The basic separation model has basic speech separation capability, making the signal distortion ratio greater than or equal to 10dB. This is only an illustrative example.
[0061] In this embodiment of the invention, a general separation model that can work in various acoustic environments is obtained through the above steps, establishing the ability to process speech signals at different distances and angles, and providing a reliable initial model for subsequent refined processing.
[0062] Step S105: Obtain multiple real scene data of the wearer and non-wearer.
[0063] In one possible implementation, obtaining multiple real scene data of the wearer and non-wearer specifically includes: obtaining multiple quiet scene speech signals of the wearer and the non-wearer; superimposing the multiple quiet scene speech signals with multiple noisy scene signals according to a set signal-to-noise ratio and signal-to-interference ratio to generate the multiple real scene data.
[0064] For example, when a wearer is reading corpus in a quiet scene after wearing the smart wearable device, the voice signal received by the smart wearable device is determined as a quiet scene voice signal; when a non-wearer is reading corpus in a quiet scene after wearing the smart wearable device, the voice signal received by the smart wearable device is determined as a quiet scene voice signal.
[0065] Step S106: Fine-tune the basic separation model based on the multiple real scene data to generate a target separation model.
[0066] In this embodiment of the invention, multiple real-world scenario data from both wearers and non-wearers are used to fine-tune the basic separation model, allowing the speech separated by the target separation model to be closer to the quality of a real person's voice. During the fine-tuning process, the basic separation model is finely adjusted using a small learning rate, which can reduce the sound quality damage and separation defects caused by inaccurate recording of the RIR function during the training of the basic separation model, thereby improving the speech quality in real-world wearing scenarios.
[0067] In one possible implementation, after step S102, the method further includes other steps, as shown in Figure 6, including the following: Step S107: Name and save the RIR function according to the set data naming format. Instruction manual, pages 7 / 11, CN 121583283 A
[0068] In one possible implementation, after determining the recording scene and recording path of the smart device, the naming format of the RIR function can be preset. For example, the naming format of the RIR function corresponding to the wearer can be rir_mouth_scene_pos.wav, where rir_mouth identifies the wearer's voice-related RIR function; scene represents the scene identifier, assuming that scene1-4 correspond to the anechoic chamber scene, conference room scene, office area scene and outdoor scene respectively; pos represents the posture identifier, assuming that pos0 is the standard posture and pos1 is the non-standard posture; the naming format of the RIR function corresponding to the wearer can be rir_talker_scene_dis*m_ang*.wav, where rir_ The talker identifier is the voice-related RIR function of the non-wearer; the scene represents the scene identifier, assuming that scene1-4 correspond to the anechoic chamber scene, conference room scene, office area scene and outdoor scene respectively; the dis*m represents the distance parameter, covering key distances such as 0.5m, 1m, 4m, etc.; the ang* represents the angle parameter, assuming that the front is 0°, the left is -90°, and the right is 90°; this is only an illustrative example and should be determined according to the actual situation.
[0069] In one possible implementation, business requirements can be analyzed according to different business scenarios, and the scene range of the wearer and / or non-wearer can be set. For example, if the business requirement is to minimize the wearer range, then the wearer area is considered as the wearer and the rest is considered as the non-wearer, which is used for business requirements such as voice calls that require maximum suppression of external interference; if the business requirement requires to customize the wearer range, for example, extracting the wearer and the front, suppressing signals in other areas, etc., it is suitable for specific business requirements, and the scene range of the wearer and / or non-wearer is determined according to the actual needs; and the multiple room pulses are obtained according to the scene range of the wearer and / or non-wearer.The response RIR function matches at least two RIR functions that are compatible with the business requirements, i.e., the RIR function of the wearer and the RIR function of the non-wearer are determined, and then targeted adaptation training is performed.
[0070] In one possible implementation, after determining the RIR function of the wearer and the non-wearer according to the business requirements, basic training data for simulation training is generated. The basic separation model is trained according to the basic training data, and then the basic separation model is fine-tuned several times according to real scene data to generate a target separation model for different business requirements.
[0071] Through the above embodiments, a large amount of open source data can be fully utilized to train the separation model, obtain effective separation effect between wearer and non-wearer, and utilize the controllability of the RIR function to obtain different separation effects for different business requirements.
[0072] In this embodiment of the invention, a target separation model is generated through a three-level progressive processing architecture of basic training, fine-tuning, and business adaptation. This makes the target separation model both general-purpose and capable of providing personalized services, thus balancing the dual requirements of algorithm performance and user experience. The differentiated RIR recording strategy for wearers and non-wearers provides a high-quality acoustic dataset specifically for smart wearable devices, significantly improving speech separation performance in complex environments. The configurable separation effect mechanism based on business needs enables personalized voice interaction experiences and supports productization requirements for different application scenarios.
[0073] In this embodiment of the invention, an apparatus for training a sound source separation model is provided, as shown in FIG7, specifically including: an acquisition unit 701, a determination unit 702, a generation unit 703, a training unit 704, and a fine-tuning unit 705; wherein, the acquisition unit 701 is used to acquire multiple sound signals from a wearer or non-wearer, wherein the multiple sound signals are generated by a smart device recording playback signals under different scenarios and different paths, and the playback signals are sound signals emitted by the wearer or non-wearer; the determination unit 702 is used to determine multiple room impulse response (RIR) functions based on the multiple sound signals; the generation unit 703 is used to generate multiple basic training data based on the multiple RIR functions; the training unit 704 is used to train a basic separation model based on the multiple basic training data; the acquisition unit 701 is also used to: acquire multiple real-world scene data of the wearer and non-wearer; the fine-tuning unit 705 is used to fine-tune the basic separation model based on the multiple real-world scene data to generate a target separation model. Instruction manual, pages 8 / 11, CN 121583283 A
[0074] Further, the determining unit is specifically used for: aligning the plurality of sound signals with the playback signal, and performing time-based adjustments on the plurality of sound signals.Inter-time correction; frequency sweeping and signal quality inspection of multiple time-corrected audio signals to determine multiple normal audio signals; deconvolution calculation of multiple normal generated audio signals to generate multiple RIR functions.
[0075] Further, the device also includes: a storage unit, used to name and save the RIR functions according to a set data naming format.
[0076] Further, the generation unit is also used to: process single-channel clean speech or multi-channel speech through the multiple RIR functions to generate basic training data, wherein the basic training data includes the input data and output data of the basic separation model.
[0077] Further, the generation unit is also used to: superimpose real noise to generate the basic training data.
[0078] Further, the basic separation model is a deep neural network model.
[0079] Further, the acquisition unit is also used to: acquire multiple quiet scene speech signals of the wearer and the non-wearer; superimpose the multiple quiet scene speech signals with multiple noise scene signals according to a set signal-to-noise ratio and signal-to-interference ratio to generate the multiple real scene data.
[0080] Further, the acquisition unit is also used to: acquire at least two RIR functions from the plurality of room impulse response RIR functions that match the business requirements, based on the scenario range of the wearer and / or non-wearer in the business requirements.
[0081] Further, the different scenarios include anechoic chamber scenarios, conference room scenarios, office area scenarios, and outdoor scenarios; the different paths include the voice path of the wearer and the voice path of the non-wearer.
[0082] FIG8 is a schematic diagram of the structure of the electronic device in an embodiment of the present invention. As shown in FIG8, it includes a general computer hardware structure, which includes at least a processor 801 and a memory 802. The processor 801 and the memory 802 are connected through a bus 803. The memory 802 is adapted to store instructions or programs executable by the processor 801. The processor 801 can be an independent microprocessor or a collection of one or more microprocessors. Thus, the processor 801 executes the instructions stored in the memory 802 to execute the method flow of the embodiment of the present invention as described above to realize the processing of data and the control of other devices. Bus 803 connects the aforementioned components together, and connects these components to display controller 804, display device, and input / output (I / O) device 805. Input / output (I / O) device 805 may be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, input / output device 805 is connected to the system via input / output (I / O) controller 806.
[0083] The instructions stored in the memory 802 are executed by at least one processor 801 to achieve the following: acquiring multiple sound signals from the wearer and / or non-wearer; determining multiple room impulse response (RIR) functions based on the multiple sound signals; generating multiple basic training data based on the multiple RIR functions; training a basic separation model based on the multiple basic training data; acquiring multiple real-world scene data from the wearer and non-wearer; and fine-tuning the basic separation model based on the multiple real-world scene data to generate a target separation model.
[0084] Specifically, the electronic device includes one or more processors 801 and a memory 802. Figure 8 uses one processor 801 as an example. The processor 801 and the memory 802 can be connected via a bus or other means. Figure 8 shows an example of connection via a bus. The memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The processor 801 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory 802, thereby realizing the above-mentioned method for determining and training the sound source separation model. Specification page 9 / 11 12 CN 121583283 A
[0085] The memory 802 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 802 may optionally include memory remotely located relative to the processor 801, and these remote memories may be connected to external devices via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0086] One or more modules are stored in the memory 802 and, when executed by one or more processors 801, perform the method of training the sound source separation model in any of the above method embodiments.
[0087] As those skilled in the art will recognize, various aspects of the embodiments of the present invention can be implemented as a system, method, or computer program product. Therefore, various aspects of the embodiments of the present invention can take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation that combines software and hardware aspects, which may generally be referred to herein as a "circuit," "module," or "system." Furthermore, various aspects of the embodiments of the present invention can take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code implemented thereon.
[0088] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example (but not limited to), an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the context of embodiments of the invention, a computer-readable storage medium can be any tangible medium capable of containing or storing a program used by or in conjunction with an instruction execution system, device, or apparatus.
[0089] A computer-readable signal medium can include a propagated digital signal having computer-readable program code implemented therein, such as in baseband or as part of a carrier wave. Such propagated signals can take any of a variety of forms, including but not limited to: electromagnetic, optical, or any suitable combination thereof. The computer-readable signal medium can be any of the following computer-readable media: not a computer-readable storage medium, and capable of communicating, propagating, or transmitting a program used by or in conjunction with an instruction execution system, device, or apparatus.
[0090] Any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, or any suitable combination thereof, can be used to transmit program code implemented on a computer-readable medium.
[0091] Computer program code for performing operations relating to aspects of embodiments of the present invention can be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++, etc.; and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can be executed entirely on a user's computer, partially on a user's computer, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server as a standalone software package. In the latter case, the remote computer can be connected to the user computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet provided by an Internet service provider).
[0092] The flowchart illustrations and / or description of the methods, apparatus (systems) and computer program products according to embodiments of the present invention are shown on pages 10 / 11 of the specification.121583283 A or a block diagram describes various aspects of embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions (executed via the processor of the computer or other programmable data processing apparatus) create means for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.
[0093] These computer program instructions can also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other apparatus to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of writing including instructions that implement the functions / actions specified in the flowchart and / or block diagram blocks or blocks.
[0094] Computer program instructions can also be loaded onto a computer, other programmable data processing equipment, or other devices to cause a series of operable steps to be performed on the computer, other programmable equipment, or other devices to produce a computer-implemented process, such that the instructions executed on the computer or other programmable equipment provide a process for implementing the function / action specified in the flowchart and / or block diagram blocks or blocks.
[0095] The above descriptions are merely preferred embodiments of this application and are not intended to limit this application. For those skilled in the art, this application can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of related data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. Users' refusal to process personal information other than the information necessary for basic functions will not affect users' use of basic functions. Instruction manual page 11 / 11, page 14, CN 121583283 A, Figure 1; Instruction manual drawing page 1 / 8, page 15, CN 121583283 A, Figure 2; Instruction manual drawing page 2 / 8, page 16, CN 121583283 A, Figure 3; Instruction manual drawing page 3 / 8, page 17, CN 121583283 A, Figure 4; Instruction manual drawing page 4 / 8, page 18, CN 121583283 A, Figure 5; Instruction manual drawing page 5 / 8, page 19, CN 121583283 A, Figure 6; Instruction manual drawing.Page 6 / 8, Figure 7 of the specification, Page 7 / 8, Figure 8 of the specification, Page 8 / 8, CN 121583283 A Abstract Abdominal ultrasound examination method, system and device separation odel The embodiment of the invention discloses a method, device and equipment for training a sound source separation model. In the embodiments of the present invention, a plurality of sound signals of a wearer and / or a non-wearer are obtained, the plurality of sound signals are generated by recording a playing signal by an intelligent device in different scenes and different paths, and the playing signal is a sound signal emitted by the wearer and / or the non-wearer; determining a plurality of room impulse response (RIR) functions according to the plurality of sound signals; generating a plurality of basic training data according to the plurality of RIR functions; training a basic separation model according to the multiple basic training data; acquiring a plurality of real scene data of a wearer and a non-wearer;and performing fine tuning on the basic separation model according to the multiple pieces of real scene data to generate a target separation model. Through the method, the target separation model is generated, the voice signal separation effect is effectively improved, and the user experience is improved.
Claims
1. A method for training a sound source separation model, characterized in that, The method includes: Acquire multiple sound signals from the wearer and / or non-wearer, wherein the multiple sound signals are generated by the smart device recording and playing signals in different scenarios and paths, and the playing signals are sound signals emitted by the wearer and / or non-wearer; Determine multiple room impulse response RIR functions based on the multiple sound signals; Multiple basic training data are generated based on the multiple RIR functions; Train a basic separation model based on the aforementioned multiple basic training data; Acquire real-world scenario data from both wearers and non-wearers; The target separation model is generated by fine-tuning the basic separation model based on the data from multiple real-world scenarios.
2. The method according to claim 1, characterized in that, The step of determining multiple room impulse response (RIR) functions based on the multiple sound signals specifically includes: Align the plurality of audio signals with the playback signal, and perform time correction on the plurality of audio signals; Multiple time-corrected audio signals were subjected to frequency sweeping and signal quality inspection to identify multiple normal audio signals. Multiple normal generated sound signals are deconvolved to generate multiple RIR functions.
3. The method according to claim 1, characterized in that, The method further includes: Name and save the RIR function according to the set data naming format.
4. The method according to claim 1, characterized in that, The generation of multiple basic training data based on the multiple RIR functions specifically includes: Single-channel clean speech or multi-channel speech is processed through the multiple RIR functions to generate basic training data, wherein the basic training data includes the input data and output data of the basic separation model.
5. The method according to claim 4, characterized in that, The method further includes: The basic training data is generated by superimposing real noise.
6. The method according to claim 1, characterized in that, The basic separation model is a deep neural network model.
7. The method according to claim 1, characterized in that, The acquisition of multiple real-world scenario data from both wearers and non-wearers specifically includes: Acquire multiple quiet scene voice signals from the wearer and the non-wearer; The multiple quiet scene voice signals and multiple noisy scene signals are superimposed according to a set signal-to-noise ratio and signal-to-scratching ratio to generate the multiple real scene data.
8. The method according to claim 1, characterized in that, The method further includes: Based on the scenario range of wearers and / or non-wearers described in the business requirements, at least two RIR functions that match the business requirements are obtained from the plurality of room impulse response RIR functions.
9. The method according to claim 1, characterized in that, The different scenarios include anechoic chamber scenarios, conference room scenarios, office area scenarios, and outdoor scenarios; the different paths include the voice path of the wearer and the voice path of the non-wearer.
10. An apparatus for training a sound source separation model, characterized in that, The device includes: The acquisition unit is used to acquire multiple sound signals from the wearer and / or non-wearer, wherein the multiple sound signals are generated by the smart device recording and playing signals in different scenarios and different paths, and the playing signals are sound signals emitted by the wearer and / or non-wearer; The determining unit is used to determine multiple room impulse response RIR functions based on the multiple sound signals; A generation unit is used to generate multiple basic training data based on the multiple RIR functions; Training unit, used to train basic separation model based on the multiple basic training data; The acquisition unit is also used to: acquire multiple real-world scenario data of the wearer and non-wearers; The fine-tuning unit is used to fine-tune the basic separation model based on the multiple real-world scenario data to generate the target separation model.
11. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-9.