Voice separation method and device, electronic equipment, storage medium and vehicle

By receiving and processing multi-channel voice signals in intelligent vehicles, extracting spatial distribution information of conversation personnel and matching with the placeholder information, the problem of mis-collection of voice interaction instructions when users are active in the car is solved, and high-quality voice interaction effects and improved user experience are achieved.

CN119943086APending Publication Date: 2025-05-06BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311465722.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In smart vehicles, when the user moves in the car, voice interaction commands are easily collected in other sound areas by mistake, causing the on-board terminal to accidentally execute commands in other sound areas, affecting the interaction effect and user experience.

Method used

By receiving the multi-channel voice signal of the current frame, sound representation information is extracted, the spatial distribution information of the conversation personnel is determined, and the placeholder information is matched to the placeholder information to separate the voice signals belonging to each placeholder area.

Benefits of technology

It effectively avoids inaccurate intention recognition and degradation of voice interaction performance caused by sound pickup caused by spatially divided sound zones, achieving good interaction effects for conversational personnel when moving in large areas of sound zones, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943086A_ABST
    Figure CN119943086A_ABST
Patent Text Reader

Abstract

The invention relates to a voice separation method and device, electronic equipment, a storage medium and a vehicle. Compared with the prior art, the voice representation extraction is carried out on the multi-channel voice signal of the current frame, the voice representation information of each session person in the current frame is determined, the spatial distribution information of each session person is further determined, the spatial distribution information of each session person is matched with the occupation information, and the space occupation information of each session person is determined. The method comprises the following steps: obtaining association information of each session person and occupation information, separating voice signals belonging to each occupation area from a multi-channel voice signal of a current frame according to the association information of each session person and occupation information, and separating the multi-channel voice signal into voice signals of each occupation area; according to the invention, inaccurate intention recognition and reduction of voice interaction performance caused by pickup of a voice area divided by space can be avoided, so that a good interaction effect can be achieved when session personnel move in a large range in the voice area, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent vehicle technology, and in particular to a speech separation method, device, electronic device, storage medium and vehicle. Background Art

[0002] With the development of intelligent technology, intelligent voice interaction technology has been widely used in various fields, especially in the automotive field. Users interact with the vehicle terminal or control the vehicle terminal to perform corresponding operations through voice interaction commands.

[0003] In the related art, the voice interaction instructions of each user are processed by dividing the sound zones. A fixed sound pickup area is divided in space as the sound zone range corresponding to each user, and a sound pickup separation algorithm is applied to pick up the sound in each sound zone to obtain the voice interaction instructions corresponding to each user.

[0004] However, since the user has a large range of activities in the car, when the user conducts voice interaction at the boundary of the sound zone or across the sound zone scenarios, it is easy for the voice interaction commands issued by the user to be collected by other sound zones, and the vehicle terminal will mistakenly think that they are voice interaction commands from other sound zones, and thus execute the voice interaction commands in the seat areas corresponding to the other sound zones. The voice interaction commands are not executed in the seat area where the user is actually located. For example, when the front passenger user is conversing with the rear seat user, the front passenger user is still sitting normally at the beginning, and the front passenger user's subsequent activity range is larger, sometimes approaching the rear sound zone, and sometimes leaning towards other sound zones. When the front passenger user's head is located on the right side of the rear row and issues a voice command to open the window, in the existing sound zone-based sound pickup solution, the vehicle terminal will mistakenly think that it is a command from the rear sound zone, and will open the rear right window, which is inconsistent with the actual command issued by the user, thereby affecting the interaction effect and the user experience. Summary of the invention

[0005] In order to solve the above technical problems, the present disclosure provides a speech separation method, device, electronic device, storage medium and vehicle.

[0006] In a first aspect, an embodiment of the present disclosure provides a speech separation method, comprising:

[0007] Receiving a multi-channel speech signal of a current frame;

[0008] Extracting sound representations of the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame;

[0009] Determine spatial distribution information of each conversation person in the current frame by combining the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames;

[0010] Matching the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain association information between each conversation person in the current frame and the placeholder information;

[0011] According to the association information between each conversation person and the placeholder information in the current frame, the voice signals belonging to each placeholder area are separated from the multi-channel voice signal of the current frame.

[0012] In some embodiments, the receiving a multi-channel speech signal of a current frame includes:

[0013] A multi-channel voice signal of a current frame collected by each sound collection device is received, wherein the multi-channel voice signal is a voice signal collected synchronously by multiple sound collection devices, and each sound collection device corresponds to one channel.

[0014] In some embodiments, extracting the sound representation of the multi-channel speech signal of the current frame to determine the sound representation information of each conversation person in the current frame includes:

[0015] The multi-channel speech signal of the current frame is input into a pre-trained sound representation extraction model, and the sound representation extraction model is used to extract the sound representation of the multi-channel speech signal of the current frame, and the sound representation information of each conversation person in the current frame is output.

[0016] In some embodiments, combining the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to determine the spatial distribution information of each conversation person in the current frame includes:

[0017] Performing sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame;

[0018] The spatial distribution information of each conversation person in the current frame is determined based on the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame. The sound source localization result of the multi-channel speech signal of each frame in the preset number of frames is obtained by performing sound source localization on the multi-channel speech signal of each frame in the preset number of frames.

[0019] In some embodiments, before determining the spatial distribution information of each conversation person in the current frame according to the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame, the method further includes:

[0020] Clustering is performed on the sound characterization information of each conversation person in each frame of the preset number of frames and the sound characterization information of each conversation person in the current frame to obtain conversation person related information of the current frame, wherein the conversation person related information of the current frame includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0021] In some embodiments, determining the spatial distribution information of each conversation person in the current frame according to the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame includes:

[0022] The class center of each conversation person in the current frame is matched with the sound source localization result of the multi-channel speech signal in the current frame to obtain the spatial distribution information of each conversation person in the current frame.

[0023] In some embodiments, matching the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain the association information between each conversation person in the current frame and the placeholder information includes:

[0024] Acquiring the vehicle's place-occupancy information, wherein the place-occupancy information includes various place-occupancy areas;

[0025] Based on the spatial distribution information of each conversation person in the current frame, obtain the spatial distribution center position of each conversation person in the current frame;

[0026] The spatial distribution center position of each conversation person in the current frame is matched with the place-occupancy information to obtain the corresponding relationship between each conversation person in the current frame and the place-occupancy information.

[0027] In some embodiments, separating the voice signals belonging to each placeholder area from the multi-channel voice signal of the current frame according to the association information between each conversation person and the placeholder information in the current frame includes:

[0028] Inputting the multi-channel speech signal of the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into a pre-trained speech separation model, performing speech separation on the multi-channel speech signal of the current frame by using the speech separation model, and outputting the speech signals belonging to each conversation person in the current frame;

[0029] According to the association information between each conversation person and the placeholder information in the current frame, the voice signals belonging to each conversation person in the current frame are matched to each placeholder area in the placeholder information associated with each conversation person to obtain the voice signals belonging to each placeholder area.

[0030] In some embodiments, the multi-channel speech signal of the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person are input into a pre-trained speech separation model, the multi-channel speech signal of the current frame is subjected to speech separation by the speech separation model, and the speech signals belonging to each conversation person in the current frame are output, including:

[0031] Inputting the multi-channel speech signal of the current frame into the encoding module of the speech separation model, encoding the multi-channel speech signal of the current frame by the encoding module to obtain an encoded multi-channel speech signal of the current frame;

[0032] Extracting the sound representation of the multi-channel speech signal of the current frame to obtain the sound representation information of each speaker in the current frame;

[0033] The encoded multi-channel speech signal of the current frame, the sound representation information of each conversation person in the current frame, and the probability that the multi-channel speech signal of the current frame belongs to each conversation person are input into the decoding module of the speech separation model, and the speech signals belonging to each conversation person and the identification of each conversation person in the current frame are output through the decoding module.

[0034] In a second aspect, an embodiment of the present disclosure provides a speech separation device, comprising:

[0035] A receiving module, used for receiving a multi-channel speech signal of a current frame;

[0036] An extraction module, configured to extract sound representation information from the multi-channel speech signal of the current frame, and determine the sound representation information of each speaker in the current frame;

[0037] a determination module, configured to determine spatial distribution information of each conversation person in the current frame by combining the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames;

[0038] An obtaining module, used for matching the spatial distribution information of each conversation person in the current frame with the placeholder information, and obtaining the association information between each conversation person in the current frame and the placeholder information;

[0039] The separation module is used to separate the voice signals belonging to each placeholder area from the multi-channel voice signal of the current frame according to the association information between each conversation person and the placeholder information in the current frame.

[0040] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:

[0041] Memory;

[0042] Processor; and

[0043] Computer programs;

[0044] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.

[0045] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method as described in the first aspect.

[0046] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the speech separation method as described above is implemented.

[0047] In a sixth aspect, an embodiment of the present disclosure further provides a vehicle, including:

[0048] Memory;

[0049] Processor; and

[0050] Computer programs;

[0051] The computer program is stored in the memory and is configured to be executed by the processor to implement the speech separation method as described above.

[0052] The speech separation method, device, electronic device, storage medium and vehicle provided by the embodiments of the present disclosure receive a multi-channel speech signal of a current frame, perform sound characterization extraction on the multi-channel speech signal of the current frame, determine the sound characterization information of each conversation person in the current frame, combine the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, determine the spatial distribution information of each conversation person in the current frame, the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel speech signal of the preset number of frames, match the spatial distribution information of each conversation person in the current frame with the placeholder information, obtain the association information of each conversation person and the placeholder information in the current frame, and separate the speech signals belonging to each placeholder area from the multi-channel speech signal of the current frame according to the association information of each conversation person and the placeholder information in the current frame. Compared with the separation scheme based on sound zone pickup in the prior art, the present invention separates based on the relevant information of the conversation persons, determines the sound representation information of multiple conversation persons in the current frame from the multi-channel voice signal of the current frame, and further combines the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame to determine the spatial distribution information of each conversation person in the current frame. The sound representation information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel voice signal of the preset number of frames. The spatial distribution information of each conversation participant in the current frame is matched with the placeholder information to obtain the association information between each conversation participant and the placeholder information in the current frame, and then according to the association information between each conversation participant and the placeholder information in the current frame, the voice signals belonging to each placeholder area are separated from the multi-channel voice signal of the current frame. Separating the multi-channel voice signal into voice signals of each placeholder area can avoid inaccurate intention recognition and decreased voice interaction performance caused by picking up sound in spatially divided sound zones, thereby achieving good interaction effects when the conversation participants move over a large range within the sound zone, thereby improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0055] Figure 1 A flow chart of a speech separation method provided in an embodiment of the present disclosure;

[0056] Figure 2 A flow chart of a speech separation method provided by another embodiment of the present disclosure;

[0057] Figure 3 A flow chart of a speech separation method provided by another embodiment of the present disclosure;

[0058] Figure 4 A schematic diagram of the structure of a pre-trained sound representation extraction model provided in an embodiment of the present disclosure;

[0059] Figure 5 A schematic diagram of the structure of a pre-trained speech separation model provided in an embodiment of the present disclosure;

[0060] Figure 6 A schematic diagram of the structure of a speech separation device provided in an embodiment of the present disclosure;

[0061] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0062] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0063] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0064] The present disclosure provides a speech separation method, which is described below in conjunction with a specific embodiment.

[0065] Figure 1A flow chart of the speech separation method provided in the embodiment of the present disclosure. The method can be applied to vehicle-mounted terminals, which can be portable mobile devices such as smart phones, tablet computers, laptop computers, vehicle-mounted navigation devices, smart sports equipment, etc.; they can also be fixed devices such as personal computers and smart home appliances. The method can be applied to scenarios where speech separation is required, and can be applied to scenarios where speech interaction occurs at the boundaries of sound zones or across sound zones, and can improve speech interaction effects and user experience. It is understandable that the speech separation method provided in the embodiment of the present disclosure can also be applied in other scenarios.

[0066] Below Figure 1 The speech separation method shown in the figure is introduced, and the specific steps of the method are as follows:

[0067] S101. Receive a multi-channel speech signal of a current frame.

[0068] In this step, the vehicle-mounted terminal receives the multi-channel voice signal of the current frame. The multi-channel voice signal is an audio signal synchronously collected by multiple sound collection devices such as microphones. The multi-channel voice signal can be a 4-channel voice signal, a 5-channel voice signal, etc. Specifically, the number of multi-channel voice signals can be set according to demand, and the position of each sound collection device in the vehicle can also be set according to demand, and this implementation does not impose any restrictions.

[0069] In some embodiments, receiving the multi-channel voice signal of the current frame includes: receiving the multi-channel voice signal of the current frame collected by each sound collection device, the multi-channel voice signal is a voice signal collected synchronously by multiple sound collection devices, and each sound collection device corresponds to one channel.

[0070] S102: Extract sound representation from the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame.

[0071] In this step, after receiving the multi-channel voice signal of the current frame, the vehicle-mounted terminal extracts the sound representation of the multi-channel voice signal of the current frame to determine the voice representation information of each conversation person in the current frame. Specifically, the multi-channel voice signal of the current frame is extracted to extract voiceprint information. Different conversation persons have different voiceprint information. Multiple conversation persons in the current frame can be determined based on the voiceprint information.

[0072] S103. Determine spatial distribution information of each conversation person in the current frame by combining the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, wherein the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel speech signal of the preset number of frames.

[0073] In this step, after determining the sound characterization information of multiple conversation persons in the current frame, the vehicle-mounted terminal determines the spatial distribution information of each conversation person in the current frame by combining the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame, wherein the sound characterization information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel voice signal of the preset number of frames. For example, there are two conversation persons, and the spatial distribution information of the two conversation persons is respectively the front row right and the back row left. In some embodiments, the spatial distribution information can be simple position indication information or the spatial coordinates of the distribution.

[0074] Exemplarily, the sound characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame can be used to identify each conversation person in the vehicle and the spatial distribution information of each conversation person in combination with the in-vehicle image of a preset number of frames before the current frame, so as to obtain the correspondence between the sound characterization information and the spatial distribution information; based on the sound characterization information of each conversation person in the current frame, the correspondence between the sound characterization information and the spatial distribution information, and the in-vehicle image of the current frame, the spatial distribution information of each conversation person in the current frame is identified.

[0075] In some embodiments, S103 may include but is not limited to S1031 and S1032:

[0076] S1031. Perform sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame.

[0077] Sound source localization is a method of estimating the location of a sound source by using the difference in strength and delay between the signals of each channel received by a microphone array. The vehicle-mounted terminal performs sound source localization on the multi-channel speech signal of the current frame, and can obtain a sound source localization result of the multi-channel speech signal of the current frame. The sound source localization result is used to characterize the location information of the sound source.

[0078] S1032. Determine spatial distribution information of each conversation person in the current frame according to a sound source localization result of the multi-channel speech signal of the current frame, a sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and sound characterization information of each conversation person in the current frame, wherein the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames is obtained by performing sound source localization on the multi-channel speech signal of each frame in the preset number of frames.

[0079] After obtaining the sound source localization result of the multi-channel voice signal of the current frame, the vehicle-mounted terminal can determine the spatial distribution information of each conversation person in the current frame based on the sound source localization result of the multi-channel voice signal of the current frame, combined with the sound source localization result of the multi-channel voice signal of each frame in the preset number of frames and the sound characterization information of each conversation person in the current frame. The sound source localization result of the multi-channel voice signal of each frame in the preset number of frames is obtained by performing sound source localization on the multi-channel voice signal of each frame in the preset number of frames. For example, if the sound source localization result of a certain frame signal is the front right of the vehicle, and one conversation person is determined in the frame, then the spatial distribution information of the conversation person is the front right. For example, if the sound source localization result of a certain frame signal is the front right and the rear right of the vehicle, and two conversation persons are determined in the frame, then the spatial distribution information of the two conversation persons is the front right and the rear right, respectively.

[0080] S104: Match the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain association information between each conversation person in the current frame and the placeholder information.

[0081] After determining the spatial distribution information of each conversation person in the current frame, the vehicle terminal can match the spatial distribution information of each conversation person in the current frame with the placeholder information in combination with the placeholder information to obtain the association information of each conversation person and the placeholder information in the current frame. For example, if the sound source localization result of a frame signal is the right front row and the right rear row of the vehicle, and two conversation persons are determined in the frame, the spatial distribution information of the two conversation persons is the right front row and the right rear row, respectively, and the placeholder information is that the right front row position and the right rear row position are occupied, then the association information of each conversation person and the placeholder information is obtained, that is, the positions corresponding to each conversation person are the right front row position and the right rear row position. For another example, the sound source localization result of a frame signal is the front left and rear right of the vehicle, and two conversation persons are determined in the frame, then the spatial distribution information of the two conversation persons are the front right and rear right respectively, and the placeholder information is that the front left position and the front right position are occupied, then the association information of each conversation person and the placeholder information is obtained, that is, the positions corresponding to each conversation person are the front left position and the front right position respectively. Although the spatial distribution information obtained based on the sound source localization result is the front right and the rear right, the actual placeholder information is that the front left position and the front right position are occupied. The spatial distribution information of each conversation person in the current frame is further matched with the placeholder information to obtain the association information of each conversation person and the placeholder information in the current frame, which can avoid inaccurate intention recognition and decreased voice interaction performance caused by picking up sound in spatially divided sound zones, thereby achieving a good interaction effect when the conversation persons move over a large range within the sound zone.

[0082] S105 . Separate speech signals belonging to respective placeholder areas from the multi-channel speech signal of the current frame according to association information between respective conversation participants and placeholder information in the current frame.

[0083] In this step, the vehicle terminal separates the voice signals belonging to each placeholder area from the multi-channel voice signal of the current frame according to the association information between each conversation person and the placeholder information in the current frame. Specifically, the vehicle terminal can obtain the voice signals of each conversation person from the multi-channel voice signal of the current frame, and combine the association information between each conversation person and the placeholder information to obtain the voice signals of each placeholder area.

[0084] Optionally, the association information between each conversation person and the placeholder information may be a correspondence between each conversation person and the placeholder information, that is, to which placeholder area in the placeholder information each conversation person corresponds.

[0085] In some embodiments, the method further includes: automatically registering the information of each conversation participant in the vehicle terminal. For example, the identity of each conversation participant, the spatial distribution information of each conversation participant, the voice characterization information of each conversation participant, the occupancy information of each conversation participant, the seat adjustment habits of each conversation participant, the listening preferences of each conversation participant and other information are automatically saved in the vehicle terminal. Through the implicit registration and matching process of the conversation participants, the cumbersome active registration process of the conversation participants is omitted, the degree of intelligence is improved, and the user experience is improved. In some optional implementations, the vehicle terminal can also monitor the changes of this information in real time, automatically update the information, and further improve the effect of voice interaction and user experience. In some embodiments, the historical related information of the conversation participants can also be saved. Combined with the historical related information, personalized interaction can be achieved, and the interaction effect can be further improved.

[0086] The embodiment of the present disclosure receives a multi-channel voice signal of a current frame, performs sound characterization extraction on the multi-channel voice signal of the current frame, determines the sound characterization information of each conversation person in the current frame, combines the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame, and determines the spatial distribution information of each conversation person in the current frame. The sound characterization information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel voice signal of the preset number of frames. The spatial distribution information of each conversation person in the current frame is matched with the placeholder information to obtain the association information between each conversation person and the placeholder information in the current frame. According to the association information between each conversation person and the placeholder information in the current frame, the voice signals belonging to each placeholder area are separated from the multi-channel voice signal of the current frame. Compared with the separation scheme based on sound zone pickup in the prior art, the present invention separates based on the relevant information of the conversation persons, determines the sound representation information of multiple conversation persons in the current frame from the multi-channel voice signal of the current frame, and further combines the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame to determine the spatial distribution information of each conversation person in the current frame. The sound representation information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel voice signal of the preset number of frames. The spatial distribution information of each conversation participant in the current frame is matched with the placeholder information to obtain the association information between each conversation participant and the placeholder information in the current frame, and then according to the association information between each conversation participant and the placeholder information in the current frame, the voice signals belonging to each placeholder area are separated from the multi-channel voice signal of the current frame. Separating the multi-channel voice signal into voice signals of each placeholder area can avoid inaccurate intention recognition and decreased voice interaction performance caused by picking up sound in spatially divided sound zones, thereby achieving good interaction effects when the conversation participants move over a large range within the sound zone, thereby improving user experience.

[0087] Figure 2 A flow chart of a speech separation method provided by another embodiment of the present disclosure is shown in FIG. Figure 2 As shown, the method includes the following steps:

[0088] S201. Receive a multi-channel speech signal of a current frame.

[0089] Specifically, the implementation process and principle of S201 and S101 are the same, and will not be repeated here.

[0090] S202: Input the multi-channel speech signal of the current frame into a pre-trained sound representation extraction model, perform sound representation extraction on the multi-channel speech signal of the current frame through the sound representation extraction model, and output sound representation information of each conversation person in the current frame.

[0091] In this step, the sound representation extraction model is pre-trained. After the vehicle terminal inputs the multi-channel speech signal of the current frame into the pre-trained sound representation extraction model, the sound representation information of each conversation person in the current frame is output through the sound representation extraction model. By establishing and training the speech separation model, the sound representation extraction of the multi-channel speech signal of the current frame is realized, thereby determining the sound representation information of each conversation person in the current frame, which facilitates the subsequent separation of the multi-channel speech signal.

[0092] Figure 4 A schematic diagram of the structure of a pre-trained sound representation extraction model provided in an embodiment of the present disclosure. Figure 4 As shown, the sound representation extraction model is composed of a speaker representation extraction module and a classifier module connected in series, and is trained with the goal of classifying each conversation person from the multi-channel speech signal of the current frame. When the training is completed, that is, the sound representation extraction model converges, the sound representation information of each conversation person in the current frame can be obtained through the speaker representation extraction module in the sound representation extraction model, and multiple conversation persons in the current frame can be obtained through the classifier in the sound representation extraction model.

[0093] In some embodiments, the sound representation extraction model may be obtained by the following steps:

[0094] 1) obtaining multiple groups of multi-channel speech signals and real extraction results corresponding to each group of multi-channel speech signals, wherein the real extraction results include multiple conversation persons and voice representation information of each conversation person;

[0095] 2) training the sound representation extraction model to be trained based on each group of multi-channel speech signals to obtain a prediction extraction result of each group of multi-channel speech signals;

[0096] 3) based on the comparison of the actual extraction results corresponding to each group of multi-channel speech signals and the predicted extraction results corresponding to each group of multi-channel speech signals, obtaining the loss calculation result of the sound representation extraction model;

[0097] 4) updating the model parameters of the sound characterization extraction model based on the loss calculation result;

[0098] 5) determining the accuracy of the updated sound representation extraction model;

[0099] 6) If the accuracy of the updated sound representation extraction model is greater than a preset threshold, the training of the sound representation extraction model is completed.

[0100] S203. Determine spatial distribution information of each conversation person in the current frame by combining the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, wherein the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel speech signal of the preset number of frames.

[0101] Specifically, the implementation process and principle of S203 and S103 are the same, and will not be repeated here.

[0102] S204: Acquire the place-occupying information of the vehicle, where the place-occupying information includes various place-occupying areas.

[0103] In this step, the vehicle terminal obtains the vehicle's occupancy information. Optionally, the vehicle's occupancy information can be obtained based on the gravity sensing data measured by the vehicle's gravity sensing device, or the vehicle's occupancy information can be identified based on the image data collected by the image acquisition device in the vehicle. For example, the front row left position and the rear row right position of the vehicle's occupancy information are obtained.

[0104] S205: Obtain the spatial distribution center position of each conversation person in the current frame based on the spatial distribution information of each conversation person in the current frame.

[0105] After determining the spatial distribution information of each conversation person in the current frame, the vehicle terminal obtains the spatial distribution center position of each conversation person in the current frame based on the spatial distribution information of each conversation person in the current frame. For example, the vehicle terminal can establish a coordinate system with the center of the vehicle as the coordinate origin, thereby obtaining the coordinates of the spatial distribution center position of each conversation person.

[0106] S206: Match the spatial distribution center position of each conversation person in the current frame with the place-occupancy information to obtain a corresponding relationship between each conversation person in the current frame and the place-occupancy information.

[0107] In this step, the vehicle-mounted terminal matches the spatial distribution center position of each conversation person in the current frame with the place-occupancy information to obtain the corresponding relationship between each conversation person in the current frame and the place-occupancy information.

[0108] For example, the spatial distribution center positions of the conversation personnel are: the spatial distribution center position of conversation personnel 1 is A, the spatial distribution center position of conversation personnel 2 is B, and the placeholder information is the front row left position and the back row right position. The spatial distribution center positions of the conversation personnel are matched with the placeholder information, and there are two matching results: A matches the front row left position, B matches the back row right position; or A matches the back row right position, B matches the front row left position. The vehicle-mounted terminal calculates the sum of the Euclidean distances between the spatial distribution center positions of the conversation personnel after matching and the center points of the placeholder areas in the placeholder information. Further, the sum of the Euclidean distances between the spatial distribution center positions of the conversation personnel and the center points of the placeholder areas in the placeholder information is calculated when A matches the front row left position and B matches the back row right position, and the sum of the Euclidean distances between the spatial distribution center positions of the conversation personnel and the center points of the placeholder areas in the placeholder information is calculated when A matches the back row right position and B matches the front row left position, and the sum of the Euclidean distances between the spatial distribution center positions of the conversation personnel and the center points of the placeholder areas in the placeholder information is calculated when A matches the back row right position and B matches the front row left position, and the sum of the Euclidean distances corresponding to various matching results are compared. When the sum of the Euclidean distances between the center positions of the spatial distribution of each conversation person and the center points of each placeholder area in the placeholder information is the smallest, it means that the center positions of the spatial distribution of each conversation person and the placeholder information are successfully matched, and the corresponding relationship between each conversation person and the placeholder information is obtained. For example, when A matches the front row left position and B matches the back row right position, the sum of the Euclidean distances is the smallest, so conversation person 1 corresponds to the front row left position and conversation person 2 corresponds to the back row right position.

[0109] S207: Separate the voice signals belonging to the respective placeholder areas from the multi-channel voice signals of the current frame according to the association information between the respective conversation participants and the placeholder information in the current frame.

[0110] Specifically, the implementation process and principle of S207 and S105 are the same, and will not be repeated here.

[0111] The embodiment of the present disclosure receives a multi-channel voice signal of the current frame, inputs the multi-channel voice signal of the current frame into a pre-trained sound representation extraction model, extracts sound representation of the multi-channel voice signal of the current frame through the sound representation extraction model, and outputs the sound representation information of each conversation person in the current frame. Further, the spatial distribution information of each conversation person in the current frame is determined by combining the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame, the sound representation information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by extracting sound representation of the multi-channel voice signal of the preset number of frames, and the occupancy information of the vehicle is obtained, the occupancy information includes each occupancy area, the spatial distribution center position of each conversation person in the current frame is obtained based on the spatial distribution information of each conversation person in the current frame, the spatial distribution center position of each conversation person in the current frame is matched with the occupancy information, and the corresponding relationship between each conversation person in the current frame and the occupancy information is obtained. Then, based on the association information between each conversation participant and the placeholder information in the current frame, the voice signals belonging to each placeholder area are separated from the multi-channel voice signal of the current frame. Compared with the separation scheme based on sound zones in the prior art, the present invention performs separation based on the relevant information of the conversation participants, determines the sound characterization information of multiple conversation participants in the current frame from the multi-channel voice signal of the current frame, and further combines the sound characterization information of each conversation participant in the current frame and the sound characterization information of each conversation participant in the multi-channel voice signal of a preset number of frames before the current frame to determine the spatial distribution information of each conversation participant in the current frame. The sound characterization information of each conversation participant in the multi-channel voice signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel voice signal of the preset number of frames. Based on the spatial distribution information of each conversation participant in the current frame, the sound characterization information of each conversation participant in the current frame is obtained. The spatial distribution center positions of the conversation participants are determined, and the spatial distribution center positions of the conversation participants are matched with the placeholder information to obtain the association information between the conversation participants and the placeholder information, and then according to the association information between the conversation participants and the placeholder information in the current frame, the voice signals belonging to the respective placeholder areas are separated from the multi-channel voice signals of the current frame. The multi-channel voice signals are separated into voice signals of the respective placeholder areas, which can avoid inaccurate intention recognition and decreased voice interaction performance caused by picking up sound in the spatially divided sound zones, thereby achieving good interaction effects when the conversation participants move over a large range within the sound zones, improving the experience of the conversation participants across sound zones and at the boundaries of sound zones, achieving high-quality interaction effects covering the entire vehicle, and improving the user experience.

[0112] Figure 3A flow chart of a speech separation method provided by another embodiment of the present disclosure is shown in FIG. Figure 3 As shown, the method includes the following steps:

[0113] S301: Receive a multi-channel speech signal of a current frame.

[0114] Specifically, the implementation process and principle of S301 and S101 are the same, and will not be repeated here.

[0115] S302: Extract sound representation from the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame.

[0116] Specifically, the implementation process and principle of S302 and S102 are the same, and will not be repeated here.

[0117] S303, clustering the sound characterization information of each conversation person in each frame of the preset number of frames and the sound characterization information of each conversation person in the current frame to obtain conversation person related information of the current frame, wherein the conversation person related information of the current frame includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0118] In this step, the vehicle-mounted terminal clusters the sound characterization information of each conversation person in each frame of the preset number of frames and the sound characterization information of each conversation person in the current frame to obtain information related to the conversation persons in the current frame. The information related to the conversation persons in the current frame includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0119] Specifically, clustering can be performed using the following formula to obtain the cluster center of each conversation participant:

[0120]

[0121] Among them, μ k is the spatial class center vector of the kth conversation person (K in total), x i represents the sound source localization result corresponding to the i-th frame, is the probability that the speech signal of the i-th frame (N frames in total) belongs to the k-th speaker and satisfies

[0122]

[0123] S304: Match the class center of each conversation person in the current frame with the sound source localization result of the multi-channel speech signal of the current frame to obtain the spatial distribution information of each conversation person in the current frame.

[0124] After obtaining the class center of each conversation person, the vehicle terminal matches the class center of each conversation person in the current frame with the sound source localization result of the multi-channel voice signal of the current frame to obtain the spatial distribution information of each conversation person in the current frame. Optionally, the class center of each conversation person is a coordinate vector, and the sound source localization result of the multi-channel voice signal of the current frame is the spatial position of the sound source of the multi-channel voice signal of the current frame. The class center of each conversation person is matched with the sound source localization result of the multi-channel voice signal of the current frame to obtain the spatial distribution information of each conversation person in the vehicle. For example, there are two conversation persons in the current frame. After matching the class centers of the two conversation persons with the sound source localization result corresponding to the frame, the spatial distribution information of the two conversation persons in the vehicle can be obtained as the front right and the back left, respectively.

[0125] S305: Match the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain association information between each conversation person in the current frame and the placeholder information.

[0126] Specifically, the implementation process and principle of S305 and S104 are the same, and will not be repeated here.

[0127] S306. Input the multi-channel speech signal of the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into a pre-trained speech separation model, perform speech separation on the multi-channel speech signal of the current frame through the speech separation model, and output the speech signals belonging to each conversation person in the current frame.

[0128] In this step, the speech separation model is pre-trained. After the vehicle terminal inputs the multi-channel speech signal of the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into the pre-trained speech separation model, the speech separation model outputs the speech signals belonging to each conversation person in the current frame. By establishing and training the speech separation model, the separation of the multi-channel speech signal of the current frame is achieved, the accuracy of speech separation is improved, and thus the speech interaction effect is improved.

[0129] The speech separation model is trained with the goal of separating the speech signals of each conversation participant from the multi-channel speech signal of the current frame. When the training is completed, that is, the speech separation model converges, the multi-channel speech signal of the current frame can be speech separated by the speech separation model, and the speech signals belonging to each conversation participant in the current frame can be output.

[0130] Figure 5 A schematic diagram of the structure of a pre-trained speech separation model provided in an embodiment of the present disclosure. Figure 5 As shown, the speech separation model includes an encoding module and a decoding module.

[0131] In some embodiments, S306 may include but is not limited to S3061, S3062, S3063:

[0132] S3061, inputting the multi-channel speech signal of the current frame into the encoding module of the speech separation model, encoding the multi-channel speech signal of the current frame by the encoding module, and obtaining an encoded multi-channel speech signal of the current frame;

[0133] S3062, extracting sound representation from the multi-channel speech signal of the current frame to obtain sound representation information of each conversation person in the current frame;

[0134] S3063. Input the encoded multi-channel speech signal of the current frame, the sound representation information of each conversation person in the current frame, and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into a decoding module of the speech separation model, and output the speech signals belonging to each conversation person and the identification of each conversation person in the current frame through the decoding module.

[0135] In some embodiments, the speech separation model may be obtained by the following steps:

[0136] 1) obtaining multiple groups of multi-channel voice signals, the probability that each group of multi-channel voice signals belongs to each conversation person, and the true separation result corresponding to each group of multi-channel voice signals, wherein the true separation result includes the voice signals belonging to each conversation person in each group of multi-channel voice signals;

[0137] 2) training a speech separation model to be trained based on each group of multi-channel speech signals to obtain a predicted separation result of each group of multi-channel speech signals;

[0138] 3) based on the comparison of the actual separation result corresponding to each group of multi-channel speech signals and the predicted separation result corresponding to each group of multi-channel speech signals, obtaining the loss calculation result of the speech separation model;

[0139] 4) based on the loss calculation result, updating the model parameters of the speech separation model;

[0140] 5) determining the accuracy of the updated speech separation model;

[0141] 6) If the accuracy of the updated speech separation model is greater than a preset threshold, the training of the speech separation model is completed.

[0142] S307: According to the association information between each conversation person and the placeholder information in the current frame, match the voice signal belonging to each conversation person in the current frame to each placeholder area in the placeholder information associated with each conversation person to obtain the voice signal belonging to each placeholder area.

[0143] In this step, the vehicle terminal matches the voice signals belonging to each conversation person in the current frame to each placeholder area in the placeholder information associated with each conversation person according to the association information between each conversation person in the current frame, and obtains the voice signals belonging to each placeholder area. For example, the vehicle terminal can match the voice signals belonging to each conversation person in the current frame to each placeholder area in combination with the association information between each conversation person and the placeholder information, and then obtain the voice signals of each placeholder area, which can avoid the situation where the instructions executed by the vehicle terminal are inconsistent with the instructions actually issued by the user, and avoid the inaccurate intention recognition and reduced voice interaction performance caused by picking up sound in the sound zones divided by space, so that the conversation person can have a good interaction effect even when moving over a large range in the sound zone, thereby improving the user experience.

[0144] The embodiment of the present disclosure receives a multi-channel speech signal of the current frame, extracts sound representation of the multi-channel speech signal of the current frame, and determines the sound representation information of each conversation person in the current frame. Further, clustering is performed on the sound representation information of each conversation person in each frame of the preset number of frames and the sound representation information of each conversation person in the current frame to obtain the conversation person related information of the current frame, the conversation person related information of the current frame includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person, the class center of each conversation person in the current frame is matched with the sound source localization result of the multi-channel speech signal of the current frame to obtain the spatial distribution information of each conversation person in the current frame, the spatial distribution information of each conversation person in the current frame is matched with the placeholder information, and the association information of each conversation person and the placeholder information in the current frame is obtained. Then, the multi-channel voice signal of the current frame and the probability that the multi-channel voice signal of the current frame belongs to each conversation person are input into a pre-trained voice separation model, and the multi-channel voice signal of the current frame is voice separated by the voice separation model, and the voice signals belonging to each conversation person in the current frame are outputted. According to the association information between each conversation person and the placeholder information in the current frame, the voice signals belonging to each conversation person in the current frame are matched to each placeholder area in the placeholder information associated with each conversation person, and the voice signals belonging to each placeholder area are obtained. This can avoid the situation where the instructions executed by the vehicle terminal are inconsistent with the instructions actually issued by the user, and avoid the inaccurate intention recognition and the degradation of voice interaction performance caused by the sound zone divided by space for sound pickup, so that the conversation person can have a good interaction effect even when moving over a large range in the sound zone, improve the experience of the conversation person across the sound zone and the sound zone boundary position, and achieve high-quality interaction effect covering the whole vehicle.

[0145] Figure 6 The structure diagram of the speech separation device provided in the embodiment of the present disclosure is shown in FIG. The speech separation device may be the vehicle-mounted terminal as described in the above embodiment, or the speech separation device may be a component or assembly in the vehicle-mounted terminal. The speech separation device provided in the embodiment of the present disclosure may execute the processing flow provided in the speech separation method embodiment, such as Figure 6As shown, the speech separation device 40 includes: a receiving module 41, an extraction module 42, a determination module 43, a obtaining module 44, and a separation module 45; wherein the receiving module 41 is used to receive a multi-channel speech signal of a current frame; the extraction module 42 is used to perform sound characterization extraction on the multi-channel speech signal of the current frame to determine the sound characterization information of each conversation person in the current frame; the determination module 43 is used to combine the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to determine the spatial distribution information of each conversation person in the current frame, wherein the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel speech signal of the preset number of frames; the obtaining module 44 is used to match the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain the association information between each conversation person and the placeholder information in the current frame; the separation module 45 is used to separate the speech signals belonging to each placeholder area from the multi-channel speech signal of the current frame according to the association information between each conversation person and the placeholder information in the current frame.

[0146] Optionally, when the receiving module 41 receives the multi-channel voice signal of the current frame, it is specifically used to: receive the multi-channel voice signal of the current frame collected by each sound collection device, the multi-channel voice signal is a voice signal collected synchronously by multiple sound collection devices, and each sound collection device corresponds to one channel.

[0147] Optionally, the extraction module 42 performs sound representation extraction on the multi-channel speech signal of the current frame, and when determining the sound representation information of each conversation participant in the current frame, it is specifically used to: input the multi-channel speech signal of the current frame into a pre-trained sound representation extraction model, perform sound representation extraction on the multi-channel speech signal of the current frame through the sound representation extraction model, and output the sound representation information of each conversation participant in the current frame.

[0148] Optionally, when the determination module 43 determines the spatial distribution information of each conversation person in the current frame by combining the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, it is specifically used to: perform sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame; determine the spatial distribution information of each conversation person in the current frame according to the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame, wherein the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames is obtained by performing sound source localization on the multi-channel speech signal of each frame in the preset number of frames.

[0149] Optionally, before determining the spatial distribution information of each conversation person in the current frame based on the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame, the speech separation device 40 also includes: a clustering module 46; the clustering module 46 is used to perform clustering processing on the sound characterization information of each conversation person in each frame in the preset number of frames and the sound characterization information of each conversation person in the current frame to obtain conversation person related information of the current frame, the conversation person related information of the current frame including the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0150] Optionally, when the determination module 43 determines the spatial distribution information of each conversation person in the current frame based on the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame, it is specifically used to: match the class center of each conversation person in the current frame with the sound source localization result of the multi-channel speech signal of the current frame to obtain the spatial distribution information of each conversation person in the current frame.

[0151] Optionally, when the obtaining module 44 matches the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain the association information between each conversation person in the current frame and the placeholder information, it is specifically used to: obtain the placeholder information of the vehicle, the placeholder information including each placeholder area; obtain the spatial distribution center position of each conversation person in the current frame based on the spatial distribution information of each conversation person in the current frame; match the spatial distribution center position of each conversation person in the current frame with the placeholder information to obtain the correspondence between each conversation person in the current frame and the placeholder information.

[0152] Optionally, when the separation module 45 separates the voice signals belonging to each placeholder area from the multi-channel voice signal of the current frame according to the association information between each conversation person and the placeholder information in the current frame, it is specifically used to: input the multi-channel voice signal of the current frame and the probability that the multi-channel voice signal of the current frame belongs to each conversation person into a pre-trained voice separation model, perform voice separation on the multi-channel voice signal of the current frame through the voice separation model, and output the voice signals belonging to each conversation person in the current frame; according to the association information between each conversation person and the placeholder information in the current frame, match the voice signals belonging to each conversation person in the current frame to each placeholder area in the placeholder information associated with the each conversation person, and obtain the voice signals belonging to each placeholder area.

[0153] Optionally, the separation module 45 inputs the multi-channel speech signal of the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into a pre-trained speech separation model, performs speech separation on the multi-channel speech signal of the current frame through the speech separation model, and outputs the speech signals belonging to each conversation person in the current frame. Specifically, it is used to: input the multi-channel speech signal of the current frame into the encoding module of the speech separation model, encode the multi-channel speech signal of the current frame through the encoding module to obtain the encoded multi-channel speech signal of the current frame; extract sound representation of the multi-channel speech signal of the current frame to obtain the sound representation information of each conversation person in the current frame; input the encoded multi-channel speech signal of the current frame, the sound representation information of each conversation person in the current frame, and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into the decoding module of the speech separation model, and output the speech signals belonging to each conversation person and the identification of each conversation person in the current frame through the decoding module.

[0154] Figure 6 The speech separation device of the illustrated embodiment can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effect are similar and will not be repeated here.

[0155] Figure 7 Schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 7 , which shows a structural schematic diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0156] like Figure 7 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503 to implement the speech separation method of the embodiment described in the present disclosure. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0157] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0158] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains a program code for executing the method shown in the flowchart, thereby implementing the speech separation method as described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0159] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0160] In addition, an embodiment of the present disclosure further provides a vehicle, comprising: a memory; a processor; and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the speech separation method as described above.

[0161] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0162] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0163] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0164] Receiving a multi-channel speech signal of a current frame;

[0165] Extracting sound representations of the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame;

[0166] Determine spatial distribution information of each conversation person in the current frame by combining the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames;

[0167] Matching the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain association information between each conversation person in the current frame and the placeholder information;

[0168] According to the association information between each conversation person and the placeholder information in the current frame, the voice signals belonging to each placeholder area are separated from the multi-channel voice signal of the current frame.

[0169] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.

[0170] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0171] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0172] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.

[0173] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0174] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0175] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0176] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0177] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

Claims

1. A speech separation method, characterized in that: The method comprises: Receiving a multi-channel speech signal of a current frame; Extracting sound representations of the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame; Determine spatial distribution information of each conversation person in the current frame by combining the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames; Matching the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain association information between each conversation person in the current frame and the placeholder information; According to the association information between each conversation person and the placeholder information in the current frame, the voice signals belonging to each placeholder area are separated from the multi-channel voice signal of the current frame.

2. The method according to claim 1, characterized in that The extracting the sound representation of the multi-channel speech signal of the current frame to determine the sound representation information of each conversation person in the current frame includes: The multi-channel speech signal of the current frame is input into a pre-trained sound representation extraction model, and the sound representation extraction model is used to extract the sound representation of the multi-channel speech signal of the current frame, and the sound representation information of each conversation person in the current frame is output.

3. The method according to claim 1, characterized in that The combining the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to determine the spatial distribution information of each conversation person in the current frame includes: Performing sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame; The spatial distribution information of each conversation person in the current frame is determined based on the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame. The sound source localization result of the multi-channel speech signal of each frame in the preset number of frames is obtained by performing sound source localization on the multi-channel speech signal of each frame in the preset number of frames.

4. The method according to claim 3, characterized in that Before determining the spatial distribution information of each conversation person in the current frame according to the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame, the method further includes: Clustering is performed on the sound characterization information of each conversation person in each frame of the preset number of frames and the sound characterization information of each conversation person in the current frame to obtain conversation person related information of the current frame, wherein the conversation person related information of the current frame includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

5. The method according to claim 4, characterized in that The determining of the spatial distribution information of each conversation person in the current frame according to the sound source localization result of the multi-channel speech signal of the current frame, the sound source localization result of the multi-channel speech signal of each frame in the preset number of frames, and the sound characterization information of each conversation person in the current frame includes: The class center of each conversation person in the current frame is matched with the sound source localization result of the multi-channel speech signal in the current frame to obtain the spatial distribution information of each conversation person in the current frame.

6. The method according to claim 4, characterized in that The step of separating the voice signals belonging to the respective placeholder areas from the multi-channel voice signals of the current frame according to the association information between the respective conversation persons and the placeholder information in the current frame includes: Inputting the multi-channel speech signal of the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into a pre-trained speech separation model, performing speech separation on the multi-channel speech signal of the current frame by using the speech separation model, and outputting the speech signals belonging to each conversation person in the current frame; According to the association information between each conversation person and the placeholder information in the current frame, the voice signals belonging to each conversation person in the current frame are matched to each placeholder area in the placeholder information associated with each conversation person to obtain the voice signals belonging to each placeholder area.

7. The method according to claim 6, characterized in that The method includes inputting the multi-channel speech signal of the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person into a pre-trained speech separation model, performing speech separation on the multi-channel speech signal of the current frame by using the speech separation model, and outputting the speech signals belonging to each conversation person in the current frame, including: Inputting the multi-channel speech signal of the current frame into the encoding module of the speech separation model, encoding the multi-channel speech signal of the current frame by the encoding module to obtain an encoded multi-channel speech signal of the current frame; Extracting the sound representation of the multi-channel speech signal of the current frame to obtain the sound representation information of each speaker in the current frame; The encoded multi-channel speech signal of the current frame, the sound representation information of each conversation person in the current frame, and the probability that the multi-channel speech signal of the current frame belongs to each conversation person are input into the decoding module of the speech separation model, and the speech signals belonging to each conversation person and the identification of each conversation person in the current frame are output through the decoding module.

8. The method according to claim 1, characterized in that The matching of the spatial distribution information of each conversation person in the current frame with the placeholder information to obtain the association information between each conversation person in the current frame and the placeholder information includes: Acquiring the vehicle's place-occupancy information, wherein the place-occupancy information includes various place-occupancy areas; Based on the spatial distribution information of each conversation person in the current frame, obtain the spatial distribution center position of each conversation person in the current frame; The spatial distribution center position of each conversation person in the current frame is matched with the place-occupancy information to obtain the corresponding relationship between each conversation person in the current frame and the place-occupancy information.

9. A speech separation device, characterized in that: include: A receiving module, used for receiving a multi-channel speech signal of a current frame; An extraction module, configured to extract sound representation information from the multi-channel speech signal of the current frame, and determine the sound representation information of each speaker in the current frame; A determination module, configured to determine spatial distribution information of each conversation person in the current frame by combining the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames; An obtaining module, used for matching the spatial distribution information of each conversation person in the current frame with the placeholder information, and obtaining the association information between each conversation person in the current frame and the placeholder information; The separation module is used to separate the voice signals belonging to each placeholder area from the multi-channel voice signal of the current frame according to the association information between each conversation person and the placeholder information in the current frame.

10. An electronic device, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A vehicle, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 8.