Voice separation method and apparatus, electronic device, and storage medium
By combining the processing of voice signals and image sequences, the problem of the inability to determine the person to which the voice signals belong in the prior art is solved, accurate control of the vehicle-mounted equipment is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202210609847.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The existing blind source separation method cannot determine the correspondence between the separated voice signal and the passengers in the vehicle, resulting in the inability to effectively manage the control of different passengers in different vehicle equipment, and the user experience is poor.
By acquiring the mixed voice signals and image sequences in the spatial area, performing image quality detection, and depending on whether the image quality meets the preset standards, different voice separation models are selected to process the mixed voice signals, separate out multiple voice signals, and determine their belongings, thereby achieving accurate control of the on-board equipment.
It realizes accurate separation and recognition of voice signals of different passengers, and can control the on-board equipment based on permission information, improving the user experience.
Smart Images

Figure CN114974245B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of vehicle technology and speech processing technology, and in particular to a speech separation method and apparatus, an electronic device, and a storage medium. Background Art
[0002] With the development of voice-based control technology and vehicle technology, a way to control in-vehicle devices by voice has emerged.
[0003] In order to facilitate multi-user control of a vehicle in the vehicle, it is necessary to separately obtain voice signals of different passengers in the vehicle. In related technologies, blind source separation (BSS) is used to separate voices. Blind source separation refers to the process of separating each source signal from a mixed signal (observed signal) when the theoretical model of the signal and the source signals cannot be accurately known. In existing blind source separation methods, since the correspondence between the separated voice signals and the passengers in the vehicle cannot be determined, effective management of controlling different in-vehicle devices by different passengers cannot be achieved. Summary of the Invention
[0004] Currently, when separating a mixed voice signal by blind source separation, voices of different persons can be obtained after separation. Since it is not clear which person the voice signal belongs to, different permission controls cannot be performed on the separated voice signals, and thus it is not clear whether to respond to the control content of the in-vehicle device for the voice signal, resulting in poor user experience.
[0005] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a speech separation method and apparatus, an electronic device, and a storage medium.
[0006] According to a first aspect of an embodiment of the present disclosure, a speech separation method is provided, including:
[0007] Obtain a first mixed voice signal and a first image sequence in a spatial region, where the first mixed voice signal includes a voice signal of a first person and a voice signal of a second person, and the first image sequence is an image sequence including the persons in the spatial region collected in the spatial region;
[0008] Perform image quality detection on the first image sequence to determine the image quality of the first image sequence;
[0009] In response to the image quality of the first image sequence meeting a preset standard, use a first speech separation model to process the input first mixed voice signal and the first image sequence to obtain a first voice signal, where the first voice signal includes at least one voice signal separated from the mixed voice signal;
[0010] In response to the image quality of the first image sequence not meeting the preset standard, the first mixed speech signal is processed using a second speech separation model to obtain a second speech signal, where the second speech signal includes at least one speech signal separated from the mixed speech signal.
[0011] According to a second aspect of the embodiments of the present disclosure, a speech separation device is provided, including:
[0012] An acquisition module, configured to acquire a first mixed speech signal and a first image sequence within a spatial region, where the first mixed speech signal includes a speech signal of a first person and a speech signal of a second person, and the first image sequence is an image sequence collected within the spatial region and including images of the people within the space;
[0013] An image quality determination unit, configured to perform image quality detection on the first image sequence to determine the image quality of the first image sequence;
[0014] A first processing unit, configured to, in response to the image quality of the first image sequence meeting the preset standard, process the input first mixed speech signal and the first image sequence using a first speech separation model to obtain a first speech signal, where the first speech signal includes at least one speech signal separated from the first mixed speech signal;
[0015] A second processing unit, configured to, in response to the image quality of the first image sequence not meeting the preset standard, process the first mixed speech signal using a second speech separation model to obtain a second speech signal, where the second speech signal includes at least one speech signal separated from the first mixed speech signal.
[0016] According to a third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, where the storage medium stores a computer program, and the computer program is used to execute the speech separation method described in the first aspect above.
[0017] According to a fourth aspect of the embodiments of the present disclosure, an electronic device is provided, where the electronic device includes:
[0018] A processor;
[0019] A memory for storing executable instructions executable by the processor;
[0020] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the speech separation method described in the first aspect above.
[0021] Based on the voice separation method, device, electronic device, and storage medium provided in the above embodiments of the present disclosure, after obtaining the first mixed voice signal and the first image sequence in a spatial region (such as a cockpit), the image quality of the first image sequence is detected to determine the image quality of the first image sequence; according to whether the image quality of the first image sequence meets a preset standard, the first voice separation model or the second voice separation model is correspondingly used to perform targeted voice separation on the first mixed voice signal, and the separated multiple voice signals can be obtained, and the person to whom the multiple voice signals belong can be determined. At least one voice signal is output from the multiple voice signals, and then, according to the permission information of the person to whom the output at least one voice signal belongs, it can be determined whether to control the in-vehicle device to respond to the voice command of the output at least one voice signal, and the user experience is good.
[0022] The technical solutions of the present disclosure will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] By describing the embodiments of the present disclosure in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present disclosure will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0024] Figure 1 is a flowchart of the voice separation method in an embodiment of the present disclosure;
[0025] Figure 2 is a flowchart of step S2 in an embodiment of the present disclosure;
[0026] Figure 3 is a flowchart of step S4 in an embodiment of the present disclosure;
[0027] Figure 4 is a structural block diagram of the voice separation device in an embodiment of the present disclosure;
[0028] Figure 5 is a structural block diagram of the image quality determination module 200 in an embodiment of the present disclosure;
[0029] Figure 6 is a block diagram of the second processing module 400 in an embodiment of the present disclosure;
[0030] Figure 7 is a structural diagram of an electronic device provided in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0032] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0033] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0034] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may mean two or more, and "at least one" may mean one, two, or more.
[0035] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, in the absence of a clear limitation or a contrary indication in the context, it can generally be understood as one or more.
[0036] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0037] It should also be understood that the present disclosure emphasizes the differences between the various embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be described one by one.
[0038] The following description of at least one exemplary embodiment is actually merely illustrative and in no way restricts the present disclosure or its application or use.
[0039] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.
[0040] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0041] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0042] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, target programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0043] Exemplary Overview
[0044] An image acquisition device and a voice acquisition device are arranged within a specified spatial region, and a mixed audio signal and an image sequence within the specified spatial region are respectively acquired through the image acquisition device and the audio acquisition device. For example, an in-vehicle camera and an in-vehicle microphone array are arranged in a cockpit, and an in-vehicle mixed audio signal and an in-vehicle image sequence are respectively acquired through the in-vehicle camera and the in-vehicle microphone array.
[0045] After obtaining the mixed audio signal within the specified spatial region, the background noise (such as wind noise or mechanical noise) in the mixed audio signal can be denoised, and then a first mixed voice signal including the voice signals of a first person and a second person can be separated from the mixed audio signal based on the audio features of the mixed audio signal.
[0046] Perform image quality detection on the first image sequence using a preset standard to determine the image quality of the first image sequence. When the image quality of the first image sequence meets the preset standard, use the pre-trained first speech separation model to process the input first mixed speech signal and the first image sequence to obtain the first speech signal; when the image quality of the first image sequence does not meet the preset standard, it is difficult to assist in speech separation of the first mixed speech signal through the first image sequence at this time. Therefore, use the second speech separation model including a blind source separation model and a person sound source model to process the first mixed speech signal to obtain the second speech signal. Among them, for the first speech signal and the second speech signal, the person to whom they belong can be determined, and then whether to respond to the control instruction of the in-vehicle device for the first speech signal or the second speech signal can be determined according to the permission information of the person to whom they belong, providing a good user experience.
[0047] Exemplary Method
[0048] Figure 1 It is a schematic flow chart of the speech separation method in an embodiment of the present disclosure. As Figure 1 shown, it includes the following steps:
[0049] S1: Obtain the first mixed speech signal and the first image sequence within the spatial region. Among them, the first mixed speech signal includes the speech signals of the first person and the second person, and may also include the speech signals of other people within the spatial region. The first image sequence is an image sequence collected within the spatial region including the people within the space.
[0050] Set up an image acquisition device and a speech acquisition device within the spatial region, and respectively collect the first mixed audio signal and the first image sequence within the specified spatial region through the image acquisition device and the audio acquisition device. Among them, the spatial region can be the cockpit, the image acquisition device can include an in-vehicle camera, and the audio acquisition device can include an in-vehicle microphone array, that is, the in-vehicle mixed audio signal and the in-vehicle image sequence can be respectively collected through the in-vehicle camera and the in-vehicle microphone array.
[0051] After obtaining the mixed audio signal within the specified spatial region, the background noise (such as wind noise or mechanical noise) in the mixed audio signal can be denoised, and then the first mixed speech signal can be separated from the mixed audio signal based on the audio features of the mixed audio signal.
[0052] S2: Perform image quality detection on the first image sequence to determine the image quality of the first image sequence.
[0053] The preset criteria may include criteria for the dimensions of the image signal or criteria for the image quality dimension. By using the preset criteria to perform image quality detection on the first image sequence, it can be determined whether the first image sequence meets the preset criteria, and then, based on the satisfaction of the first image sequence with respect to the preset criteria, the image quality of the first image sequence can be determined.
[0054] S3: In response to the image quality of the first image sequence meeting the preset criteria, use the first voice separation model to process the input first mixed voice signal and the first image sequence to obtain a first voice signal. Wherein, the first voice signal includes at least one voice signal separated from the first mixed voice signal.
[0055] When the image quality of the first image sequence meets the preset criteria, use the first mixed voice signal and the first image sequence as the input of the pre-trained first voice separation model, and use the first voice separation model to process the first mixed voice signal and the first image sequence to obtain a first voice signal.
[0056] Wherein, before step S3, it is possible to collect, within a predetermined time period, a mixed voice signal including multiple persons for a spatial region as a sample mixed voice signal. Collect an image sequence including these multiple persons for the spatial region within the predetermined time period as a sample image sequence. Based on the sample mixed voice signal and the sample image sequence, train the first voice separation model. Wherein, the first voice separation model can perform image recognition on the sample image sequence to determine the speaking time and speaking content of these multiple persons within the predetermined time period. The first voice separation model can also perform voice separation on the sample voice signal based on the voice features of different voice signals in the sample mixed voice signal to obtain multiple voice signals. Furthermore, the first voice separation model can determine the sound source object of the multiple voice signals based on the speaking time and speaking content of these multiple persons within the predetermined time period, so as to determine the persons to whom the multiple voice signals belong.
[0057] S4: In response to the image quality of the first image sequence not meeting the preset criteria, use the second voice separation model to process the first mixed voice signal to obtain a second voice signal, where the second voice signal includes at least one voice signal separated from the first mixed voice signal.
[0058] When the image quality of the first image sequence does not meet the preset standard, it is difficult to determine the person to whom the multiple voice signals after separating the first mixed voice signal can be assisted by the image recognition result of the first image sequence. Therefore, the second voice separation model can be used to process the first mixed voice signal to obtain the second voice signal. Among them, the second voice separation model can include a blind source separation model and a sound source model for determining the sound source object of the multiple voice signals. The blind source separation model can be used to separate the first mixed voice signal into multiple voice signals, and then the sound source model can be used to determine the person to whom the multiple voice signals belong, and then the first voice signal can be output, and the person to whom the first voice signal belongs can also be output.
[0059] In this embodiment, after obtaining the first mixed voice signal and the first image sequence in the spatial region, the image quality of the first image sequence is detected to determine the image quality of the first image sequence. According to whether the image quality of the first image sequence meets the preset standard, the first voice separation model or the second voice separation model is correspondingly used to perform targeted voice separation on the first mixed voice signal, and the separated multiple voice signals can be obtained. At least one voice signal is output from the multiple voice signals, and then, according to the permission information of the person to whom at least one voice signal is output, it can be determined whether to control the vehicle-mounted device to respond to the voice command of the output at least one voice signal, and the user experience is good.
[0060] Figure 2 It is a schematic flowchart of step S2 in an embodiment of the present disclosure. As Figure 2 shown, step S2 includes:
[0061] S2-1: Obtain the image signal corresponding to the first image sequence and determine the image signal quality of the image signal.
[0062] The signal intensity of the image signal corresponding to the first image sequence can be detected, and the image signal quality can be determined according to the comparison result between the detected signal intensity and the preset signal intensity threshold. For example, when the detected signal intensity is greater than the signal intensity threshold, it is determined that the image signal intensity meets the image signal quality standard; when the detected signal intensity is less than or equal to the signal intensity threshold, it is determined that the image signal intensity does not meet the image signal quality standard.
[0063] S2-2: Based on each image frame of the first image sequence, determine the image content quality of the first image sequence.
[0064] Image recognition can be performed on each image frame of the first image sequence, and the image content quality of the first image sequence can be determined based on the image recognition result. For example, when the image recognition result can assist in determining the person to whom the speech signal belongs after separating the first mixed speech signal, it is determined that the image content quality meets the preset image content quality standard; when the image recognition result cannot assist in determining the person to whom the speech signal belongs after separating the first mixed speech signal, it is determined that the image content quality does not meet the preset image content quality standard.
[0065] S2-3: Determine the image quality of the first image sequence based on the image signal quality and the image content quality.
[0066] It can be set that when any one of the image signal quality and the image content quality does not meet the corresponding quality standard, it is determined that the image quality of the first image sequence does not meet the preset standard. For example, when the image signal quality does not meet the image signal quality standard, steps S2-2 to S2-3 are not entered, and it is directly determined that the image quality of the first image sequence does not meet the preset standard.
[0067] In this embodiment, by detecting the first image sequence, the image signal quality and the image content quality of the first image sequence can be determined. Through the two detection dimension qualities of the image signal quality and the image content quality, the image quality of the first image sequence can be effectively characterized.
[0068] In an embodiment of the present disclosure, step S2-2 specifically includes: determining the lip occlusion state of the first person and / or the second person in each image frame based on each image frame of the first image sequence; determining the image content quality based on the lip occlusion state.
[0069] The lip occlusion state may include: the lips are not occluded and the lips are occluded. When the lips are occluded, it is impossible to determine whether the first person and / or the second person is speaking through the lip image, and it is even more impossible to determine the speaking time of the first person and / or the second person. Furthermore, it is impossible to determine the person to whom the multiple speech signals belong after separating the first mixed speech signal based on the speaking time.
[0070] When the lip occlusion state of the first person and / or the second person in each image frame is that the lips are occluded, it is determined that the image content quality of the first image sequence does not meet the image content quality standard.
[0071] When the lip occlusion state of the first person and / or the second person in each image frame is that the lips are not occluded, the lip image of the first person and / or the second person in each image frame can be obtained and recognized based on the lip image. If the recognition result of the lip image can determine the speaking time and speaking content of the first person and / or the second person, it can be determined that the image content quality of the first image sequence meets the image content quality standard.
[0072] In this embodiment, based on each image frame of the first image sequence, the lip occlusion state of the first person and / or the second person in each image frame can be determined. Based on the lip occlusion state, it can be quickly determined whether the image content quality of the first image sequence meets the image content quality standard, and then it can be quickly determined whether the image quality of the first image sequence meets the preset standard, so that it can be quickly determined whether to select the first voice separation model or the second voice separation model for voice separation.
[0073] In an embodiment of the present disclosure, step S2-3 includes: in response to the image signal quality not meeting the image signal quality standard, determining that the image quality of the first image sequence does not meet the preset standard; in response to the image content quality not meeting the image content quality standard, determining that the image quality of the first image sequence does not meet the preset standard; in response to the image signal quality meeting the image signal quality standard and the image content quality meeting the image content quality standard, determining that the image quality of the first image sequence meets the preset standard.
[0074] In this embodiment, when the image signal quality does not meet the image signal quality standard, the clarity of the first image sequence usually generated based on the image signal is insufficient, and it is difficult to analyze the person to whom the multiple voice signals after separating the first mixed voice signal belong based on the first image sequence. When the image content quality does not meet the image content quality standard, the speaking time and speaking content of the first person and / or the second person cannot be obtained, and thus the person to whom the multiple voice signals after separating the first mixed voice signal belong cannot be determined. Therefore, only when both the image signal quality and the image quality meet the corresponding quality standards can it be determined that the image quality of the first image sequence meets the preset standard, and then based on the recognition result of the first image sequence, the person to whom the multiple voice signals after separating the first mixed voice signal belong can be effectively determined. When any one of the image signal quality and the image quality does not meet the corresponding quality standard, it can be quickly determined that the image quality of the first image sequence does not meet the preset standard.
[0075] In an embodiment of the present disclosure, step S2-2 includes: in response to the lip occlusion state being that the lips of the first person and / or the second person are not occluded, based on each image frame of the first image sequence, determining the lip movements of the first person and / or the second person; in response to the lip movements not meeting the preset lip movement standard, determining that the image quality of the first image sequence does not meet the image content quality standard.
[0076] If the lips of the first person and / or the second person are not occluded in each image frame of the first image sequence, a lip image block sequence of the first person and / or the second person in the first image sequence can be obtained, and based on the recognition of the lip image block sequence, the lip movements of the first person and / or the second person can be obtained.
[0077] Obtain a preset lip movement standard corresponding to lip-reading recognition. Among them, the lip movement standard can include, for example, that the upper and lower rows of teeth do not touch when a person is speaking. Such a setting can filter out lip movement behaviors of eating food. Through the preliminary movement standard, lip movement behaviors other than those of a person speaking can be filtered out.
[0078] When the lip movement does not meet the preset lip movement standard, it is difficult to perform effective lip-reading recognition based on the first image sequence, and thus it is impossible to accurately determine the person to whom the multiple voice signals after separating the first mixed voice signal belong.
[0079] In this embodiment, when the lips of the first person and / or the second person in each image frame of the first image sequence are not blocked, the lip movement of the first person and / or the second person can be determined based on each image frame of the first image sequence. When the lip movement does not meet the preset lip movement standard, it indicates that it is difficult to perform effective lip-reading recognition based on the first image sequence, and thus it is impossible to accurately determine the person to whom the multiple voice signals after separating the first mixed voice signal belong.
[0080] In an embodiment of the present disclosure, the second voice separation model includes a first person sound source model, a second person sound source model, and a blind source separation model. Among them, the blind source separation model is used to perform blind source separation on the first mixed voice signal, the first person sound source model is used to determine the voice signal of the first person based on the result of the blind source separation, and the second person sound source model is used to determine the voice signal of the second person based on the result of the blind source separation.
[0081] When using the second voice separation model to process the first mixed voice signal in step S4, the blind source separation model is used to perform voice separation on the first mixed voice signal to obtain multiple voice signals. Among them, the voice signal of each person corresponds to one voice signal, and the multiple voice signals at least include the voice signal of the first person and the voice signal of the second person. When the first mixed voice signal further includes the voice signals of other persons, the multiple voice signals further include the voice signals of other persons.
[0082] After obtaining the multiple voice signals, the multiple voice signals are respectively subjected to sound source feature matching with the first person sound source model and the second person sound source model, so as to determine which voice signal has the first person as the sound source object and which voice signal has the second person as the sound source object.
[0083] After determining the voice signal corresponding to the first person and the voice signal corresponding to the second person, the second voice signal is determined based on the voice signal corresponding to the first person and / or the voice signal corresponding to the second person, and then the second voice signal is output.
[0084] In this embodiment, when the image quality of the first image sequence does not meet the preset standard, it is difficult to assist in determining the person to whom the multiple voice signals after separating the first mixed voice signal are attributed based on the image recognition result of the first image sequence. Therefore, a blind source separation model can be used to perform voice separation on the first mixed voice signal to obtain multiple voice signals, and the first person sound source model and the second person sound source model are used to perform sound source feature matching with the multiple voice signals obtained by the blind source separation model, so as to determine which voice signal has the first person as the sound source object and which voice signal has the second person as the sound source object, thereby enabling accurate voice signal separation and sound source object matching of the first mixed voice signal.
[0085] In an embodiment of the present disclosure, before using the second voice separation model to process the first mixed voice signal to obtain the second voice signal, it further includes:
[0086] A: Process the second mixed voice signal and the second image sequence based on the first voice separation model to obtain the sound source signal of the first person and the sound source signal of the second person. Among them, the acquisition time of the second mixed voice signal is earlier than the acquisition time of the first mixed voice signal. The acquisition time of the second image sequence is earlier than the acquisition time of the first image sequence. The second mixed voice signal includes the sound source signal of the first person and the sound source signal of the second person. The second image sequence is an image sequence collected in a spatial region and including the people in the space.
[0087] Since the first voice separation model can perform voice separation by combining voice channel separation and lip reading recognition, the first voice separation model can be used to accurately separate the voice and match the sound source object of the second mixed voice signal, so as to obtain the sound source signal of the first person and the sound source signal of the second person.
[0088] B: Perform online modeling based on the sound source signal of the first person and the sound source signal of the second person to obtain the first person sound source model and the second person sound source model.
[0089] The sound source features can be extracted from the sound source signal of the first person, and then the first person sound source model can be trained by the online modeling scoring method. And the sound source features can be extracted from the sound source signal of the second person, and then the second person sound source model can be trained by the online modeling scoring method.
[0090] In this embodiment, the first voice separation model is used to accurately separate the voice and match the sound source object of the second mixed voice signal, so as to obtain the sound source signal of the first person and the sound source signal of the second person. Online modeling is performed based on the sound source signal of the first person and the sound source signal of the second person, and the first person sound source model and the second person sound source model can be quickly obtained.
[0091] Figure 3 It is a schematic flowchart of step S4 in an embodiment of the present disclosure. As Figure 3 shown, step S4 includes:
[0092] S4-1: Based on the first image sequence, determine the identity information of the first person and / or the second person.
[0093] The face features of the first person and the second person can be pre-stored, and the identity information of the first person and / or the second person can be determined by means of face feature matching.
[0094] S4-2: Based on the identity information, obtain the first person sound source model and / or the second person sound source model that matches the identity information, and obtain the blind source separation model.
[0095] After determining the identity information of the first person and / or the second person, since the first person sound source model and the second person sound source model are pre-trained, the first person sound source model and / or the second person sound source model in the second voice separation model can be retrieved, and the blind source separation model can be retrieved.
[0096] S4-3: Process the first mixed voice signal based on the first person sound source model and / or the second person sound source model, and the blind source separation model to obtain the second voice signal.
[0097] Use the blind source separation model to perform voice separation on the first mixed voice signal to obtain multiple voice signals, and use the first person sound source model and the second person sound source model to perform sound source feature matching with the multiple voice signals obtained by the blind source separation model, so as to determine which voice signal's sound source object is the first person and which voice signal's sound source object is the second person, and then obtain the second voice signal.
[0098] In this embodiment, the identity information of the first person and / or the second person can be recognized from the first image sequence, and then the first person sound source model and / or the second person sound source model can be retrieved, and the blind source separation model can be retrieved, so that when the image quality of the first image sequence does not meet the preset standard, the blind source separation model can be combined with the sound source object model to accurately perform voice separation on the first mixed voice signal.
[0099] Any voice separation method provided by the embodiments of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers, etc. Or, any voice separation method provided by the embodiments of the present disclosure can be executed by a processor. For example, the processor executes any voice separation method mentioned in the embodiments of the present disclosure by calling the corresponding instructions stored in the memory. This will not be elaborated below.
[0100] Exemplary Device
[0101] Figure 4 It is a structural block diagram of a voice separation device in an embodiment of the present disclosure. As Figure 4 shown, the voice separation device includes:
[0102] An acquisition module 100, configured to acquire a first mixed voice signal and a first image sequence within a spatial region, where the first mixed voice signal includes a voice signal of a first person and a voice signal of a second person, and the first image sequence is an image sequence including people in the space collected within the spatial region;
[0103] An image quality determination module 200, configured to perform image quality detection on the first image sequence to determine the image quality of the first image sequence;
[0104] A first processing module 300, configured to, in response to the image quality of the first image sequence meeting a preset standard, process the input first mixed voice signal and the first image sequence using a first voice separation model to obtain a first voice signal, where the first voice signal includes at least one voice signal separated from the first mixed voice signal;
[0105] A second processing module 400, configured to, in response to the image quality of the first image sequence not meeting the preset standard, process the first mixed voice signal using a second voice separation model to obtain a second voice signal, where the second voice signal includes at least one voice signal separated from the first mixed voice signal.
[0106] Figure 5 It is a structural block diagram of the image quality determination module 200 in an embodiment of the present disclosure. As Figure 5 shown, the image quality determination module 200 includes:
[0107] An image signal quality determination unit 210, configured to acquire an image signal corresponding to the first image sequence and determine the image signal quality of the image signal;
[0108] An image content quality determination unit 220, configured to determine the image content quality of the first image sequence based on each image frame of the first image sequence;
[0109] An image quality determination unit 230, configured to determine the image quality of the first image sequence based on the image signal quality and the image content quality.
[0110] In one embodiment of the present disclosure, the image content quality determination unit 220 is configured to determine the lip occlusion state of the first person and / or the second person in each image frame based on each image frame of the first image sequence; the image content quality determination unit 220 is further configured to determine the image content quality based on the lip occlusion state.
[0111] In one embodiment of the present disclosure, the image content quality determination unit 220 is configured to determine that the image quality of the first image sequence does not meet the preset standard in response to the image signal quality not meeting the image signal quality standard; the image content quality determination unit 220 is further configured to determine that the image quality of the first image sequence does not meet the preset standard in response to the image content quality not meeting the image content quality standard; the image content quality determination unit 220 is further configured to determine that the image quality of the first image sequence meets the preset standard in response to the image signal quality meeting the image signal quality standard and the image content quality meeting the image content quality standard.
[0112] In one embodiment of the present disclosure, the image content quality determination unit 220 is configured to determine the lip movement of the first person and / or the second person based on each image frame of the first image sequence in response to the lip occlusion state being that the lips of the first person and / or the second person are not occluded; the image content quality determination unit 220 is further configured to determine that the image quality of the first image sequence does not meet the image content quality standard in response to the lip movement not meeting the preset lip movement standard.
[0113] In one embodiment of the present disclosure, the second voice separation model includes a first person sound source model, a second person sound source model, and a blind source separation model, wherein the blind source separation model is configured to perform blind source separation on the first mixed voice signal, the first person sound source model is configured to determine the voice signal of the first person based on the result of the blind source separation, and the second person sound source model is configured to determine the voice signal of the second person based on the result of the blind source separation.
[0114] In one embodiment of the present disclosure, the second processing module 400 is specifically configured to process the second mixed speech signal and the second image sequence based on the first speech separation model to obtain the sound source signal of the first person and the sound source signal of the second person, where the acquisition time of the second mixed speech signal is earlier than the acquisition time of the first mixed speech signal, the acquisition time of the second image sequence is earlier than the acquisition time of the first image sequence, the second mixed speech signal includes the sound source signal of the first person and the sound source signal of the second person, and the second image sequence is an image sequence including people in the space region collected in the space region; the second processing module 400 is further configured to perform online modeling based on the sound source signal of the first person and the sound source signal of the second person to obtain the first person sound source model and the second person sound source model.
[0115] Figure 6 is a block diagram of the second processing module 400 in one embodiment of the present disclosure. As Figure 6 shown, the second processing module 400 includes:
[0116] An identity information determination module 410, configured to determine the identity information of the first person and / or the second person based on the first image sequence;
[0117] A model acquisition unit 420, configured to acquire the first person sound source model and / or the second person sound source model that match the identity information based on the identity information, and acquire the blind source separation model;
[0118] A voice signal determination unit 430, configured to process the first mixed speech signal based on the first person sound source model and / or the second person sound source model, and the blind source separation model to obtain the second speech signal.
[0119] It should be noted that the specific implementation manner of the speech separation device in the embodiment of the present disclosure is similar to the specific implementation manner of the speech separation method in the embodiment of the present disclosure. For details, please refer to the speech separation method section. To reduce redundancy, it will not be elaborated here.
[0120] Exemplary Electronic Device
[0121] Next, refer to Figure 7 to describe the electronic device according to the embodiment of the present disclosure. As Figure 7 shown, the electronic device includes one or more processors 10 and a memory 20.
[0122] The processor 10 may be a central processing unit (CPU) or other form of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0123] The memory 20 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 11 may run the program instructions to implement the voice separation method of each embodiment of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage media.
[0124] In one example, the electronic device may further include: an input device 30 and an output device 40, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 30 may be, for example, a keyboard, a mouse, etc. The output device 40 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0125] Of course, for simplicity, Figure 7 only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0126] Exemplary Computer - Readable Storage Medium
[0127] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage media include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0128] The basic principles of the present disclosure have been described in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for illustrative and facilitating understanding purposes and are not limitations. The above details do not limit the present disclosure to necessarily adopting the above specific details for implementation.
[0129] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For system embodiments, since they basically correspond to method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments.
[0130] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with each other.
[0131] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented through software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is only for illustration purposes. The steps of the methods of the present disclosure are not limited to the above specifically described order, unless otherwise specifically stated. Additionally, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers the recording medium storing the programs for executing the methods according to the present disclosure.
[0132] It also needs to be pointed out that in the apparatuses, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0133] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0134] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A voice separation method, comprising: Obtaining a first mixed voice signal and a first image sequence within a spatial region, wherein the first mixed voice signal includes the voice signal of a first person and the voice signal of a second person, and the first image sequence is an image sequence collected within the spatial region and including the people within the space; Performing image quality detection on the first image sequence to determine the image quality of the first image sequence; In response to the image quality of the first image sequence meeting a preset standard, using a first voice separation model to process the input first mixed voice signal and the first image sequence to obtain a first voice signal, wherein the first voice signal includes at least one voice signal separated from the mixed voice signal; In response to the image quality of the first image sequence not meeting the preset standard, using a second voice separation model to process the first mixed voice signal to obtain a second voice signal, wherein the second voice signal includes at least one voice signal separated from the mixed voice signal.
2. The voice separation method according to claim 1, wherein, The performing image quality detection on the first image sequence to determine the image quality of the first image sequence includes: Obtaining the image signal corresponding to the first image sequence and determining the image signal quality of the image signal; Based on each image frame of the first image sequence, determining the image content quality of the first image sequence; Based on the image signal quality and the image content quality, determining the image quality of the first image sequence.
3. The method according to claim 2, wherein The based on each image frame of the first image sequence, determining the image content quality of the first image sequence includes: Based on each image frame of the first image sequence, determining the lip occlusion state of the first person and / or the second person in each image frame; Based on the lip occlusion state, determining the image content quality.
4. The method according to claim 3, wherein, The based on the image signal quality and the image content quality, determining the image quality of the first image sequence includes: In response to the image signal quality not meeting the image signal quality standard, determining that the image quality of the first image sequence does not meet the preset standard; In response to the image content quality not meeting the image content quality standard, determining that the image quality of the first image sequence does not meet the preset standard; In response to the image signal quality meeting the image signal quality standard and the image content quality meeting the image content quality standard, determining that the image quality of the first image sequence meets the preset standard.
5. According to the method of claim 4, the based on each image frame of the first image sequence, determining the image content quality of the first image sequence further includes: In response to the lip occlusion state being that the lips of the first person and / or the second person are not occluded, based on each image frame of the first image sequence, determining the lip movement of the first person and / or the second person; In response to the lip movement not meeting the preset lip movement standard, determining that the image quality of the first image sequence does not meet the image content quality standard.
6. The voice separation method according to claim 1, wherein the second voice separation model includes a first person sound source model, a second person sound source model, and a blind source separation model, where The blind source separation model is used to perform blind source separation on the first mixed speech signal. The first person sound source model is used to determine the speech signal of the first person based on the result of the blind source separation. The second person sound source model is used to determine the speech signal of the second person based on the result of the blind source separation.
7. The speech separation method according to claim 6, before processing the first mixed speech signal by using the second speech separation model to obtain a second speech signal, further comprising: Performing online modeling based on the historical sound source signals of the first person and the second person separated by the first speech separation model to obtain the first person sound source model and the second person sound source model.
8. The voice separation method according to claim 7, wherein, The processing the first mixed speech signal by using the second speech separation model to obtain a second speech signal includes: Determining the identity information of the first person and / or the second person based on the first image sequence; Based on the identity information, obtaining the first person sound source model and / or the second person sound source model that matches the identity information, and obtaining the blind source separation model; Processing the first mixed speech signal based on the first person sound source model and / or the second person sound source model, and the blind source separation model to obtain the second speech signal.
9. A speech separation device, comprising: An acquisition module, configured to acquire a first mixed speech signal and a first image sequence in a spatial region, where the first mixed speech signal includes the speech signals of a first person and a second person, and the first image sequence is an image sequence collected in the spatial region and including the people in the space; An image quality determination module, configured to perform image quality detection on the first image sequence to determine the image quality of the first image sequence; A first processing module, configured to, in response to the image quality of the first image sequence meeting a preset standard, process the input first mixed speech signal and the first image sequence by using a first speech separation model to obtain a first speech signal, where the first speech signal includes at least one path of speech signal separated from the first mixed speech signal; A second processing module, configured to, in response to the image quality of the first image sequence not meeting the preset standard, process the first mixed speech signal by using a second speech separation model to obtain a second speech signal, where the second speech signal includes at least one path of speech signal separated from the first mixed speech signal.
10. A computer-readable storage medium, where the storage medium stores a computer program, and the computer program is used to execute the speech separation method according to any one of claims 1-8 above.
11. An electronic device, the electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the speech separation method according to any one of claims 1-8 above.
Citation Information
Patent Citations
User identification based on voice and face
US10178301B1
Lip-reading recognition method and apparatus, computer device, and storage medium
WO2020232867A1