Voice registration method and device, electronic equipment, storage medium and vehicle

By receiving and analyzing multi-channel voice signals, the voiceprint information and spatial location of the conversation personnel are automatically registered, and the cumbersome registration process in the prior art is solved, which improves the flexibility and user experience of voice registration.

CN119943056APending Publication Date: 2025-05-06BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311466304.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, users need to manually enter registration information, resulting in cumbersome voice registration process, poor flexibility and poor user experience.

Method used

By receiving multi-channel voice signals, extracting sound characterization information, performing clustering processing and probability analysis, the voiceprint information and spatial location of the conversation personnel are automatically registered.

Benefits of technology

No need for active registration for users, simplifying the registration process and improving the flexibility and user experience of voice registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943056A_ABST
    Figure CN119943056A_ABST
Patent Text Reader

Abstract

The invention relates to a voice registration method and device, electronic equipment, a storage medium and a vehicle. Compared with the prior art, sound representation extraction is performed on the multi-channel voice signal of the current frame, so that the sound representation information of each session person in the current frame is determined; performing clustering processing on the sound representation information of each session person in the current frame and the sound representation information of each session person in the multi-channel voice signals of a preset number of frames before the current frame to obtain a representation clustering result of the session person in the current frame, the voiceprint information and the spatial position of each session person are further determined, the voiceprint information and the spatial position of each session person can be automatically registered, the registration information of each session person is obtained, a user does not need to actively carry out personalized voice registration, the problem that the registration process is tedious in the prior art is solved, and the user experience is improved. The flexibility of voice registration is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent vehicle technology, and in particular to a voice registration method, device, electronic device, storage medium and vehicle. Background Art

[0002] With the development of intelligent technology, intelligent voice interaction technology has been widely used in various fields, especially in the automotive field. Users interact with the vehicle terminal or control the vehicle terminal to perform corresponding operations through voice interaction commands.

[0003] In related technologies, in order to achieve personalized voice interaction, users need to actively register on the vehicle terminal and enter relevant registration information. Only then can the vehicle terminal match the user registration information with the actual passenger information to provide a personalized interactive experience.

[0004] However, in the prior art, when users actively perform personalized voice registration, they need to manually input relevant registration information, and the registration process is relatively cumbersome, resulting in poor flexibility of voice registration and poor user experience. Summary of the invention

[0005] In order to solve the above technical problems, the present disclosure provides a voice registration method, device, electronic device, storage medium and vehicle.

[0006] In a first aspect, an embodiment of the present disclosure provides a voice registration method, including:

[0007] Receiving a multi-channel speech signal of a current frame;

[0008] Extracting sound representations of the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame;

[0009] Clustering the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain a representation clustering result of the conversation persons in the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames;

[0010] Probabilistic analysis is performed based on the representation clustering results to obtain the voiceprint information and spatial position of each conversation person;

[0011] The voiceprint information and spatial position of each conversation person are registered to obtain registration information of each conversation person.

[0012] In some embodiments, the probability analysis based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person includes:

[0013] Performing sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame;

[0014] Probabilistic analysis is performed in combination with the sound source localization result and the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person.

[0015] In some embodiments, the characterization clustering result includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person;

[0016] The clustering process is performed on the voice characterization information of each conversation person in the current frame and the voice characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain the characterization clustering result of the conversation persons in the current frame, including:

[0017] Clustering the voice characterization information of each conversation person in the current frame and the voice characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain the cluster center of each conversation person in the current frame;

[0018] Based on the sound representation information of each conversation person in the current frame, combined with the class center, class weight and covariance matrix of each conversation person, the probability that the multi-channel speech signal of the current frame belongs to each conversation person is calculated.

[0019] In some embodiments, the probability analysis is performed in combination with the sound source localization result and the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, including:

[0020] Determine the voiceprint information of each conversation person according to the cluster center of each conversation person in the characterization clustering result and the probability that the multi-channel speech signal of the current frame belongs to each conversation person;

[0021] Performing probability statistics calculation on the probability that the multi-channel speech signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel speech signal of the current frame to obtain the spatial distribution information of each conversation person;

[0022] The spatial distribution information of each conversation participant is matched with the placeholder information to determine the spatial position of each conversation participant.

[0023] In some embodiments, matching the spatial distribution information of each conversation participant with the placeholder information to determine the spatial position of each conversation participant includes:

[0024] Acquiring the vehicle's place-occupancy information, wherein the place-occupancy information includes various place-occupancy areas;

[0025] Based on the spatial distribution information of each conversation person, obtain the spatial distribution center position of each conversation person;

[0026] The spatial distribution center position of each conversation person is matched with the placeholder information to obtain the spatial position of each conversation person.

[0027] In some embodiments, matching the spatial distribution center position of each conversation person with the placeholder information to obtain the spatial position of each conversation person includes:

[0028] Matching the spatial distribution center position of each conversation person with the place-occupancy information, and calculating the sum of the Euclidean distances between the matched spatial distribution center position of each conversation person and the center point of each place-occupancy area in the place-occupancy information;

[0029] When the sum of the Euclidean distances between the spatial distribution center positions of the respective conversation participants and the center points of the respective placeholder areas in the placeholder information is the smallest, the spatial distribution center positions of the respective conversation participants are matched with the placeholder information to obtain the spatial positions of the respective conversation participants.

[0030] In some embodiments, after performing probability analysis based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, the method further includes:

[0031] Calculating the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint database;

[0032] For the voiceprint information of any conversation person, if the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint database is greater than a preset threshold, it is determined that the conversation person and the conversation person corresponding to the voiceprint information registered in the voiceprint database are the same conversation person, and the conversation person is associated with the conversation person corresponding to the voiceprint information registered in the voiceprint database;

[0033] The voiceprint information and spatial position of each conversation person are registered to obtain registration information of each conversation person, including:

[0034] The voiceprint information and spatial position of each conversation person are added to the registration information of the conversation person corresponding to the voiceprint information registered in the voiceprint library.

[0035] In some embodiments, after registering the voiceprint information and spatial position of each conversation person to obtain the registration information of each conversation person, the method further includes:

[0036] Real-time monitoring of the voiceprint information and spatial position of each registered conversation person to see if there is any change;

[0037] If the voiceprint information of each conversation person changes, or the spatial position of each conversation person changes, the registration information of each conversation person is updated.

[0038] In a second aspect, an embodiment of the present disclosure provides a voice registration device, including:

[0039] A receiving module, used for receiving a multi-channel speech signal of a current frame;

[0040] A first determination module is used to extract sound representation information from the multi-channel speech signal of the current frame to determine the sound representation information of each conversation person in the current frame;

[0041] A first obtaining module is used to perform clustering processing on the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, to obtain a representation clustering result of the conversation persons in the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames;

[0042] The second obtaining module is used to perform probability analysis based on the characterization clustering results to obtain the voiceprint information and spatial position of each conversation person;

[0043] The registration module is used to register the voiceprint information and spatial position of each conversation person to obtain the registration information of each conversation person.

[0044] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:

[0045] Memory;

[0046] Processor; and

[0047] Computer programs;

[0048] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.

[0049] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method as described in the first aspect.

[0050] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the voice registration method as described above is implemented.

[0051] In a sixth aspect, an embodiment of the present disclosure further provides a vehicle, including:

[0052] Memory;

[0053] Processor; and

[0054] Computer programs;

[0055] The computer program is stored in the memory and is configured to be executed by the processor to implement the voice registration method as described above.

[0056] The voice registration method, device, electronic device, storage medium and vehicle provided by the embodiments of the present disclosure receive a multi-channel voice signal of a current frame, perform sound characterization extraction on the multi-channel voice signal of the current frame, determine the voice characterization information of each conversation person in the current frame, cluster the voice characterization information of each conversation person in the current frame and the voice characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame, obtain the characterization clustering result of the conversation persons in the current frame, perform probability analysis based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, register the voiceprint information and spatial position of each conversation person, and obtain the registration information of each conversation person. Compared with the prior art, the present invention does not require users to actively perform personalized voice registration, thereby solving the problem of the complicated registration process in the prior art. By performing sound characterization extraction on the multi-channel voice signal of the current frame, the sound characterization information of each conversation person in the current frame is determined, and the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame are clustered to obtain the characterization clustering result of the conversation persons in the current frame, and the voiceprint information and spatial position of each conversation person are further determined, and then the voiceprint information and spatial position of each conversation person can be automatically registered to obtain the registration information of each conversation person, thereby improving the flexibility of voice registration and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0059] Figure 1 A flow chart of a voice registration method provided by an embodiment of the present disclosure;

[0060] Figure 2 A flow chart of a voice registration method provided by another embodiment of the present disclosure;

[0061] Figure 3 A flow chart of a voice registration method provided by another embodiment of the present disclosure;

[0062] Figure 4 A schematic diagram of the voice registration process provided by an embodiment of the present disclosure;

[0063] Figure 5 A schematic diagram of the structure of a voice registration device provided in an embodiment of the present disclosure;

[0064] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0065] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0066] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0067] The present disclosure provides a voice registration method, which is described below in conjunction with a specific embodiment.

[0068] Figure 1A flow chart of the voice registration method provided in the embodiment of the present disclosure. The method can be applied to vehicle-mounted terminals, which can be portable mobile devices such as smart phones, tablet computers, laptops, vehicle-mounted navigation devices, smart sports equipment, etc.; they can also be fixed devices such as personal computers and smart home appliances. The method can be applied to scenarios where voices are registered, without the need for users to actively perform personalized voice registration, solving the problem of the cumbersome registration process in the prior art. The voiceprint information and spatial position of each conversation person can be automatically registered to obtain the registration information of each conversation person, thereby increasing the flexibility of voice registration and improving the user experience. It can be understood that the voice registration method provided in the embodiment of the present disclosure can also be applied in other scenarios.

[0069] Below Figure 1 The voice registration method shown in the figure is introduced, and the specific steps of the method are as follows:

[0070] S101. Receive a multi-channel speech signal of a current frame.

[0071] In this step, the vehicle-mounted terminal receives the multi-channel voice signal of the current frame. The multi-channel voice signal is an audio signal synchronously collected by multiple sound collection devices such as microphones. The multi-channel voice signal can be a 4-channel voice signal, a 5-channel voice signal, etc. Specifically, the number of multi-channel voice signals can be set according to demand, and the position of each sound collection device in the vehicle can also be set according to demand, and this implementation does not impose any restrictions.

[0072] In some embodiments, receiving the multi-channel voice signal of the current frame includes: receiving the multi-channel voice signal of the current frame collected by each sound collection device, the multi-channel voice signal is a voice signal collected synchronously by multiple sound collection devices, and each sound collection device corresponds to one channel.

[0073] S102: Extract sound representation from the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame.

[0074] In this step, after receiving the multi-channel voice signal of the current frame, the vehicle-mounted terminal extracts the sound representation of the multi-channel voice signal of the current frame to determine the voice representation information of each conversation person in the current frame. Specifically, the multi-channel voice signal of the current frame is extracted to extract the voiceprint information. Different conversation persons have different voiceprint information, and multiple conversation persons in the current frame can be determined based on the voiceprint information.

[0075] S103, clustering the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain a characterization clustering result of the conversation persons in the current frame, wherein the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel speech signal of the preset number of frames.

[0076] In this step, the vehicle terminal performs clustering processing on the voice characterization information of each conversation person in the current frame and the voice characterization information of each conversation person in the multi-channel voice signal of the preset number of frames before the current frame, and obtains the characterization clustering result of the conversation person in the current frame. The voice characterization information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by performing voice characterization extraction on the multi-channel voice signal of the preset number of frames.

[0077] In some embodiments, the characterization clustering result includes the cluster center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0078] S104: Perform probability analysis based on the characterization clustering results to obtain the voiceprint information and spatial position of each conversation participant.

[0079] After obtaining the characterization clustering result of the multi-channel speech signal of the current frame, the vehicle terminal can perform probability analysis based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person. It is understandable that the voiceprint information of different conversation persons is different.

[0080] Exemplarily, the characterization clustering results of the multi-channel speech signals of a preset number of frames before the current frame can be used in combination with the in-vehicle images of a preset number of frames before the current frame to identify the individual conversation persons in the vehicle and the spatial positions of the individual conversation persons, and obtain the corresponding relationship between the characterization clustering results and the spatial positions; based on the characterization clustering results of the individual conversation persons in the current frame, the corresponding relationship between the characterization clustering results and the spatial positions, and the in-vehicle images of the current frame, the spatial positions of the individual conversation persons in the current frame are identified.

[0081] In some embodiments, S104 includes but is not limited to S1041 and S1042:

[0082] S1041, performing sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame;

[0083] Sound source localization is a method of estimating the location of a sound source by using the difference in the intensity and delay between the signals of each channel received by the microphone array. The vehicle-mounted terminal performs sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame. The sound source localization result is used to characterize the location information of the sound source. Figure 4 As shown, the vehicle-mounted terminal performs sound source localization on the received multi-channel speech signal to obtain the sound source localization result of the multi-channel speech signal of the current frame.

[0084] S1042: Perform probability analysis on the sound source localization result and the characterization clustering result to obtain the voiceprint information and spatial position of each conversation participant.

[0085] Furthermore, the vehicle-mounted terminal may perform probability analysis on the sound source localization result and the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person.

[0086] S105: Register the voiceprint information and spatial position of each conversation participant to obtain registration information of each conversation participant.

[0087] In this step, the vehicle terminal registers the voiceprint information and spatial position of each conversation person to obtain the registration information of each conversation person. Specifically, the vehicle terminal can associate the voiceprint information and spatial position of each conversation person with the actual conversation person, register it in the vehicle terminal, and obtain the registration information of each conversation person. In some embodiments, after obtaining the registration information, it will be determined whether each conversation person in the registration information has registered the registration information. If the registration information has been registered, it will be determined whether the registration information of each conversation person has changed. If it has changed, the registered registration information will be updated according to the registration information; if the target conversation person has not registered the registration information, the registration information of the target conversation person will be stored.

[0088] In some embodiments, the method further includes: automatically registering other relevant information of each conversation person in the vehicle terminal. For example, the identity of each conversation person, the spatial distribution information of each conversation person, the voice characterization information of each conversation person, the occupancy information of each conversation person, the seat adjustment habits of each conversation person, the listening preferences of each conversation person and other relevant information are automatically saved in the vehicle terminal, and through the implicit conversation person registration and matching process, the tedious conversation person active registration process is omitted, the intelligence level is improved, and the user experience is improved.

[0089] In some embodiments, after registration is completed, the association between voice interaction and conversation person information is realized, so that functions such as conversation person-related voice separation or target conversation person extraction can be performed. At the same time, because of specific tags, the user's historical usage habits, behavior patterns and interactive context connections can be combined to achieve a more personalized interactive experience. For example, you can use a specific nickname to respond, or provide exclusive temperature, seat and other adjustment modes, eliminating tedious repetitive settings.

[0090] The disclosed embodiment receives a multi-channel voice signal of a current frame, extracts sound representation of the multi-channel voice signal of the current frame, determines the sound representation information of each conversation person in the current frame, clusters the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame, obtains a representation clustering result of the conversation persons in the current frame, performs probability analysis based on the representation clustering result, obtains voiceprint information and spatial position of each conversation person, registers the voiceprint information and spatial position of each conversation person, and obtains registration information of each conversation person. Compared with the prior art, the present invention does not require users to actively perform personalized voice registration, thereby solving the problem of the complicated registration process in the prior art. By performing sound characterization extraction on the multi-channel voice signal of the current frame, the sound characterization information of each conversation person in the current frame is determined, and the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame are clustered to obtain the characterization clustering result of the conversation persons in the current frame, and the voiceprint information and spatial position of each conversation person are further determined, and then the voiceprint information and spatial position of each conversation person can be automatically registered to obtain the registration information of each conversation person, thereby improving the flexibility of voice registration and enhancing the user experience.

[0091] On the basis of the above embodiment, after registering the voiceprint information and spatial position of each conversation person and obtaining the registration information of each conversation person, the method further includes: step A and step B:

[0092] Step A: Real-time monitoring of the voiceprint information and spatial position of each registered conversation person to see if there is any change;

[0093] After obtaining the registration information of each conversation person, the vehicle terminal will monitor in real time whether the voiceprint information and spatial position of each registered conversation person have changed.

[0094] Step B: if the voiceprint information of each conversation person changes, or the spatial position of each conversation person changes, then the registration information of each conversation person is updated.

[0095] When the voiceprint information of each conversation person changes, or the spatial position of each conversation person changes, the vehicle-mounted terminal updates the registration information of each conversation person.

[0096] The disclosed embodiment monitors the changes of such information in real time and automatically updates the information, further improving the effect of voice interaction and user experience. In some embodiments, the historical information of the conversation personnel can also be saved, and combined with the historical information, personalized interaction can be achieved to further improve the interaction effect.

[0097] Figure 2 A flow chart of a voice registration method provided by another embodiment of the present disclosure is shown in FIG. Figure 2 As shown, the method includes the following steps:

[0098] S201. Receive a multi-channel speech signal of a current frame.

[0099] Specifically, the implementation process and principle of S201 and S101 are the same, and will not be repeated here.

[0100] S202: Extract sound representation from the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame.

[0101] Specifically, the implementation process and principle of S202 and S102 are the same, and will not be repeated here.

[0102] S203, clustering the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain a characterization clustering result of the conversation persons in the current frame, wherein the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel speech signal of the preset number of frames.

[0103] Specifically, the implementation process and principle of S203 and S103 are the same, and will not be repeated here.

[0104] In some embodiments, the characterization clustering result includes the cluster center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0105] In some embodiments, a Gaussian Mixture Model (GMM) may be used to perform clustering using an EM algorithm to obtain the cluster center of each conversation person and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0106] Specifically, when clustering is performed using a Gaussian Mixture (GMM) model and an EM algorithm, S203 may include but is not limited to S2031 and S2032:

[0107] S2031, clustering the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain the cluster center of each conversation person in the current frame.

[0108] S2032. Based on the sound representation information of each conversational person in the current frame, combined with the class center, class weight and covariance matrix of each conversational person, calculate the probability that the multi-channel speech signal of the current frame belongs to each conversational person.

[0109] The probability that the multi-channel speech signal of the current frame belongs to each conversation person is:

[0110]

[0111] Among them, P(c k |x i ) is the probability that the i-th frame belongs to the k-th conversation person, x i is the voice representation information of the conversation person in the i-th frame, π k , μ k ,Σ k are the class weight, class center and covariance matrix of the kth session person respectively.

[0112] In some embodiments, clustering can also be performed using the following formula to obtain the cluster center of each conversation person:

[0113]

[0114] Among them, μ k is the spatial class center representation vector of the kth conversation person (K in total), x i represents the sound source localization result corresponding to the i-th frame, is the probability that the speech signal of the i-th frame (N frames in total) belongs to the k-th speaker, and satisfies

[0115]

[0116] S204: Perform sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame.

[0117] Sound source localization is a method of estimating the location of a sound source by using the difference in the intensity and delay between the signals of each channel received by the microphone array. The vehicle-mounted terminal performs sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame. The sound source localization result is used to characterize the location information of the sound source. Figure 4 As shown, the vehicle-mounted terminal performs sound source localization on the received multi-channel speech signal to obtain the sound source localization result of the multi-channel speech signal of the current frame.

[0118] S205: Determine the voiceprint information of each conversation person according to the cluster center of each conversation person in the characterization clustering result and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0119] In this step, the vehicle terminal determines the voiceprint information of each conversation person according to the class center of each conversation person in the characterization clustering result and the probability that the multi-channel voice signal of the current frame belongs to each conversation person. Specifically, the target conversation person to which the multi-channel voice signal of the current frame belongs is determined based on the probability that the multi-channel voice signal of the current frame belongs to each conversation person, and the class center of the target conversation person is obtained in combination with the class centers of each conversation person, and the class center of the target conversation person is determined as the voiceprint information of each conversation person.

[0120] S206: Perform probability statistics calculation on the probability that the multi-channel speech signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel speech signal of the current frame to obtain spatial distribution information of each conversation person.

[0121] In this step, the vehicle terminal can perform probability statistics calculation on the probability that the multi-channel voice signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel voice signal of the current frame to obtain the spatial distribution information of each conversation person. For example, the sound source localization result of a frame signal is the front right of the vehicle, and a conversation person is determined based on the class center of the frame, then the spatial distribution information of the conversation person is the front right. For example, the sound source localization result of a frame signal is the front right and the back right of the vehicle, and two conversation persons are determined based on the class center of the frame, then the spatial distribution information of the two conversation persons is the front right and the back right respectively. In some embodiments, the spatial distribution information can be simple position indication information or the spatial coordinates of the distribution.

[0122] S207: Match the spatial distribution information of each conversation participant with the placeholder information to determine the spatial position of each conversation participant.

[0123] After determining the spatial distribution information of each conversation person, the vehicle terminal can match the spatial distribution information of each conversation person with the placeholder information in combination with the placeholder information to determine the spatial position of each conversation person. For example, if the sound source localization result of a frame signal is the right front row and the right rear row of the vehicle, and two conversation persons are determined in the frame, then the spatial distribution information of the two conversation persons is the right front row and the right rear row, respectively, and the placeholder information is that the right front row position and the right rear row position are occupied, then the association information of each conversation person and the placeholder information is obtained, that is, the positions corresponding to each conversation person are the right front row position and the right rear row position, respectively. For another example, the sound source localization result of a frame signal is the front left and the rear right of the vehicle, and two conversation persons are identified in the frame, then the spatial distribution information of the two conversation persons is the front right and the rear right respectively, and the placeholder information is that the front left and the front right are occupied, then the association information of each conversation person and the placeholder information is obtained, that is, the corresponding positions of each conversation person are the front left and the front right respectively, although the spatial distribution information obtained based on the sound source localization result is the front right and the rear right, but the actual placeholder information is that the front left and the front right are occupied, then the spatial distribution information of each conversation person is further matched with the placeholder information to more accurately determine the spatial position of each conversation person.

[0124] S208: Register the voiceprint information and spatial position of each conversation participant to obtain registration information of each conversation participant.

[0125] Specifically, the implementation process and principle of S208 and S103 are the same, and will not be repeated here.

[0126] The embodiment of the present disclosure receives a multi-channel voice signal of the current frame, extracts sound representation of the multi-channel voice signal of the current frame, determines the sound representation information of each conversation person in the current frame, clusters the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame, and obtains the representation clustering result of the conversation person in the current frame. The sound representation information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by extracting sound representation of the multi-channel voice signal of the preset number of frames. Further, according to the class center of each conversation person in the representation clustering result and the probability that the multi-channel voice signal of the current frame belongs to each conversation person, the voiceprint information of each conversation person is determined, the sound source of the multi-channel voice signal of the current frame is localized, and the sound source localization result of the multi-channel voice signal of the current frame is obtained. The probability that the multi-channel voice signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel voice signal of the current frame are subjected to probability statistical calculation to obtain the spatial distribution information of each conversation person, and the spatial distribution information of each conversation person is matched with the occupancy information to determine the spatial position of each conversation person. Then, the voiceprint information and spatial position of each conversation person are registered to obtain the registration information of each conversation person. Compared with the prior art, the present disclosure does not require the user to actively perform personalized voice registration, which solves the problem that the registration process in the prior art is relatively cumbersome. The spatial distribution information of each conversation person is matched with the placeholder information to more accurately determine the spatial position of each conversation person, and then the voiceprint information and spatial position of each conversation person can be automatically registered to obtain the registration information of each conversation person, thereby improving the flexibility of voice registration and enhancing the user experience.

[0127] Figure 3 A flow chart of a voice registration method provided by another embodiment of the present disclosure is shown in FIG. Figure 3 As shown, the method includes the following steps:

[0128] S301: Obtain a representation clustering result of a multi-channel speech signal of a current frame.

[0129] Specifically, the vehicle terminal receives the multi-channel voice signal of the current frame, performs sound characterization extraction on the multi-channel voice signal of the current frame, determines the sound characterization information of each conversation person in the current frame, clusters the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame, and obtains the characterization clustering result of the conversation persons in the current frame, wherein the sound characterization information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel voice signal of the preset number of frames. In some embodiments, the characterization clustering result includes the class center of each conversation person in the current frame and the probability that the multi-channel voice signal of the current frame belongs to each conversation person.

[0130] S302: Determine the voiceprint information of each conversation person according to the cluster center of each conversation person in the characterization clustering result and the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0131] Specifically, the implementation process and principle of S302 and S205 are the same, and will not be repeated here.

[0132] S303: Perform probability statistics calculation on the probability that the multi-channel speech signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel speech signal of the current frame to obtain spatial distribution information of each conversation person.

[0133] Specifically, the implementation process and principle of S303 and S206 are the same, and will not be repeated here.

[0134] S304: Acquire the place-occupying information of the vehicle, where the place-occupying information includes various place-occupying areas.

[0135] In this step, the vehicle terminal obtains the vehicle's occupancy information. Optionally, the vehicle's occupancy information can be obtained based on the gravity sensing data measured by the vehicle's gravity sensing device, or the vehicle's occupancy information can be identified based on the image data collected by the image acquisition device in the vehicle. For example, the front row left position and the rear row right position of the vehicle's occupancy information are obtained.

[0136] S305: Obtain the spatial distribution center position of each conversation participant based on the spatial distribution information of each conversation participant.

[0137] After determining the spatial distribution information of each conversation person, the vehicle terminal obtains the spatial distribution center position of each conversation person based on the spatial distribution information of each conversation person. For example, the vehicle terminal can establish a coordinate system with the center of the vehicle as the coordinate origin, thereby obtaining the coordinates of the spatial distribution center position of each conversation person.

[0138] S306: Match the spatial distribution center position of each conversation participant with the placeholder information to obtain the spatial position of each conversation participant.

[0139] In this step, the vehicle-mounted terminal matches the spatial distribution center position of each conversation person with the occupancy information to obtain the spatial position of each conversation person.

[0140] In some embodiments, S306 includes but is not limited to S3061 and S3062:

[0141] S3061. Match the spatial distribution center positions of the respective conversation participants with the place-occupying information, and calculate the sum of the Euclidean distances between the matched spatial distribution center positions of the respective conversation participants and the center points of the respective place-occupying areas in the place-occupying information.

[0142] For example, the spatial distribution center positions of the conversation persons are: the spatial distribution center position of conversation person 1 is A, the spatial distribution center position of conversation person 2 is B, and the placeholder information is respectively the front row left position and the back row right position. The spatial distribution center positions of the conversation persons are matched with the placeholder information, and there are two matching results: A matches the front row left position, B matches the back row right position; or A matches the back row right position, B matches the front row left position. The vehicle-mounted terminal calculates the sum of the Euclidean distances between the spatial distribution center positions of the conversation persons after matching and the center points of the placeholder areas in the placeholder information.

[0143] S3062: When the sum of the Euclidean distances between the spatial distribution center positions of the respective conversation participants and the center points of the respective placeholder areas in the placeholder information is the smallest, the spatial distribution center positions of the respective conversation participants are matched with the placeholder information to obtain the spatial positions of the respective conversation participants.

[0144] As in the example described in S3061, in this step, the sum of the Euclidean distances between the center positions of the spatial distribution of each conversation person and the center points of each placeholder area in the placeholder information when A matches the left position in the front row and B matches the right position in the back row is calculated, and the sum of the Euclidean distances between the center positions of the spatial distribution of each conversation person and the center points of each placeholder area in the placeholder information when A matches the right position in the back row and B matches the left position in the front row is calculated, and the sums of the Euclidean distances corresponding to various matching results are compared. When the sum of the Euclidean distances between the center positions of the spatial distribution of each conversation person and the center points of each placeholder area in the placeholder information is the smallest, it means that the center positions of the spatial distribution of each conversation person successfully match the placeholder information, and the spatial positions of each conversation person are obtained. For example, when A matches the left position in the front row and B matches the right position in the back row, the sum of the Euclidean distances is the smallest, then conversation person 1 corresponds to the left position in the front row and conversation person 2 corresponds to the right position in the back row.

[0145] S307: Calculate the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint database.

[0146] The essence of this step is to match the voiceprint information of each conversation person with the voiceprint information registered in the voiceprint database. Specifically, the vehicle terminal calculates the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint database. Figure 4 As shown, the vehicle-mounted terminal calculates the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint database.

[0147] For example, the cosine similarity formula is used to calculate the similarity. The cosine similarity formula is:

[0148]

[0149] in, is the similarity between the current k-th conversation person and the p-th conversation person in the voiceprint database, μ k , They are the class center representation vectors of the current k-th conversation person and the p-th conversation person in the voiceprint library respectively.

[0150] S308. For the voiceprint information of any conversation person, if the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint library is greater than a preset threshold, it is determined that the conversation person and the conversation person corresponding to the voiceprint information registered in the voiceprint library are the same conversation person, and the conversation person is associated with the conversation person corresponding to the voiceprint information registered in the voiceprint library.

[0151] For the voiceprint information of any conversation person, if it is determined that the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint library is greater than a preset threshold, it is determined that the conversation person and the conversation person corresponding to the voiceprint information registered in the voiceprint library are the same conversation person, and the conversation person is associated with the conversation person corresponding to the voiceprint information registered in the voiceprint library. In some embodiments, if it is determined that the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint library is less than or equal to a preset threshold, it is determined that the conversation person is a new conversation person, that is, a conversation person who has not registered.

[0152] S309: Add the voiceprint information and spatial position of each conversation person to the registration information of the conversation person corresponding to the voiceprint information registered in the voiceprint database.

[0153] After associating the conversation personnel with the conversation personnel corresponding to the voiceprint information registered in the voiceprint library, further, the voiceprint information and spatial position of each conversation personnel are added to the registration information of the conversation personnel corresponding to the voiceprint information registered in the voiceprint library.

[0154] The disclosed embodiment obtains the characterization clustering result of the multi-channel speech signal of the current frame, determines the voiceprint information of each conversation person according to the class center of each conversation person in the characterization clustering result and the probability that the multi-channel speech signal of the current frame belongs to each conversation person, performs probability statistics calculation on the probability that the multi-channel speech signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel speech signal of the current frame, and obtains the spatial distribution information of each conversation person. Further, the vehicle's occupancy information is obtained, the occupancy information includes each occupancy area, the spatial distribution center position of each conversation person is obtained based on the spatial distribution information of each conversation person, and the spatial distribution center position of each conversation person is matched with the occupancy information to obtain the spatial position of each conversation person. Then, the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint library is calculated. For any conversation person's voiceprint information, if the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint library is greater than a preset threshold, it is determined that the conversation person and the conversation person corresponding to the voiceprint information registered in the voiceprint library are the same conversation person, and the conversation person is associated with the conversation person corresponding to the voiceprint information registered in the voiceprint library. Then, the voiceprint information and spatial position of each conversation person are added to the registration information of the conversation person corresponding to the voiceprint information registered in the voiceprint library. The present disclosure does not require users to actively perform personalized voice registration, which solves the problem of the complicated registration process in the prior art. By calculating the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint library, for the voiceprint information of any conversation person, if the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint library is greater than a preset threshold, it is determined that the conversation person and the conversation person corresponding to the voiceprint information registered in the voiceprint library are the same conversation person, and the conversation person is associated with the conversation person corresponding to the voiceprint information registered in the voiceprint library. Then, the voiceprint information and spatial position of each conversation person are added to the registration information of the conversation person corresponding to the voiceprint information registered in the voiceprint library, thereby improving the flexibility of voice registration and enhancing user experience.

[0155] Figure 5 The structure diagram of the voice registration device provided in the embodiment of the present disclosure is shown in FIG. The voice registration device may be the vehicle-mounted terminal as described in the above embodiment, or the voice registration device may be a component or assembly in the vehicle-mounted terminal. The voice registration device provided in the embodiment of the present disclosure may execute the processing flow provided in the voice registration method embodiment, such as Figure 5As shown, the voice registration device 50 includes: a receiving module 51, a first determining module 52, a first obtaining module 53, a second obtaining module 54, and a registration module 55; wherein the receiving module 51 is used to receive a multi-channel voice signal of a current frame; the first determining module 52 is used to perform sound characterization extraction on the multi-channel voice signal of the current frame to determine the voice characterization information of each conversation person in the current frame; the first obtaining module 53 is used to perform clustering processing on the voice characterization information of each conversation person in the current frame and the voice characterization information of each conversation person in the multi-channel voice signal of a preset number of frames before the current frame to obtain the characterization clustering result of the conversation person in the current frame, wherein the voice characterization information of each conversation person in the multi-channel voice signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel voice signal of the preset number of frames; the second obtaining module 54 is used to perform probability analysis based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person; the registration module 55 is used to register the voiceprint information and spatial position of each conversation person to obtain the registration information of each conversation person.

[0156] Optionally, when the second obtaining module 54 performs probability analysis based on the representation clustering result to obtain the voiceprint information and spatial position of each conversation person, it is specifically used to: perform sound source positioning on the multi-channel speech signal of the current frame to obtain the sound source positioning result of the multi-channel speech signal of the current frame; and perform probability analysis on the sound source positioning result and the representation clustering result to obtain the voiceprint information and spatial position of each conversation person.

[0157] Optionally, the characterization clustering result includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person; the first obtaining module 53 clusters the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain the characterization clustering result of the conversation persons in the current frame, and when the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound characterization extraction on the multi-channel speech signal of the preset number of frames, it is specifically used to: cluster the sound characterization information of each conversation person in the current frame and the sound characterization information of each conversation person in the multi-channel speech signal of the preset number of frames before the current frame to obtain the class center of each conversation person in the current frame; based on the sound characterization information of each conversation person in the current frame, combined with the class center, class weight and covariance matrix of each conversation person, calculate the probability that the multi-channel speech signal of the current frame belongs to each conversation person.

[0158] Optionally, when the second obtaining module 54 performs probability analysis in combination with the sound source localization result and the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, it is specifically used to: determine the voiceprint information of each conversation person based on the class center of each conversation person in the characterization clustering result and the probability that the multi-channel voice signal of the current frame belongs to each conversation person; perform probability statistical calculation on the probability that the multi-channel voice signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel voice signal of the current frame to obtain the spatial distribution information of each conversation person; match the spatial distribution information of each conversation person with the occupancy information to determine the spatial position of each conversation person.

[0159] Optionally, the second obtaining module 54 matches the spatial distribution information of each conversation participant with the placeholder information to determine the spatial position of each conversation participant, and is specifically used to: obtain the placeholder information of the vehicle, the placeholder information including each placeholder area; obtain the central position of the spatial distribution of each conversation participant based on the spatial distribution information of each conversation participant; match the central position of the spatial distribution of each conversation participant with the placeholder information to obtain the spatial position of each conversation participant.

[0160] Optionally, when the second obtaining module 54 matches the spatial distribution center position of each conversation participant with the placeholder information to obtain the spatial position of each conversation participant, it is specifically used to: match the spatial distribution center position of each conversation participant with the placeholder information, and calculate the sum of the Euclidean distances between the matched spatial distribution center position of each conversation participant and the center point of each placeholder area in the placeholder information; when the sum of the Euclidean distances between the spatial distribution center position of each conversation participant and the center point of each placeholder area in the placeholder information is the smallest, the spatial distribution center position of each conversation participant is matched with the placeholder information to obtain the spatial position of each conversation participant.

[0161] Optionally, after the probability analysis is performed based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, the device 50 further includes: a calculation module 56 and a second determination module 57; the calculation module 56 is used to calculate the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint library; the second determination module 57 is used to determine that the conversation person and the conversation person corresponding to the voiceprint information registered in the voiceprint library are the same conversation person if the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint library is greater than a preset threshold, and associate the conversation person with the conversation person corresponding to the voiceprint information registered in the voiceprint library;

[0162] The registration module 55 registers the voiceprint information and spatial position of each conversation person, and when obtaining the registration information of each conversation person, is specifically used to: add the voiceprint information and spatial position of each conversation person to the registration information of the conversation person corresponding to the voiceprint information registered in the voiceprint library.

[0163] Optionally, after registering the voiceprint information and spatial position of each conversation person and obtaining the registration information of each conversation person, the device further includes: an updating module 58; the updating module 58 is used to monitor in real time whether the voiceprint information and spatial position of each registered conversation person have changed; if the voiceprint information of each conversation person has changed, or the spatial position of each conversation person has changed, the registration information of each conversation person is updated.

[0164] Figure 5 The voice registration device of the illustrated embodiment can be used to implement the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effect are similar and will not be repeated here.

[0165] Figure 6 Schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 6 , which shows a structural schematic diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0166] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603 to implement the voice registration method of the embodiment described in the present disclosure. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0167] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0168] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains a program code for executing the method shown in the flowchart, thereby implementing the voice registration method as described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0169] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0170] In addition, an embodiment of the present disclosure further provides a vehicle, comprising: a memory; a processor; and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the voice registration method as described above.

[0171] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0172] The computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.

[0173] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0174] Receiving a multi-channel speech signal of a current frame;

[0175] Extracting sound representations of the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame;

[0176] Clustering the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain a representation clustering result of the conversation persons in the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames;

[0177] Probabilistic analysis is performed based on the representation clustering results to obtain the voiceprint information and spatial position of each conversation person;

[0178] The voiceprint information and spatial position of each conversation person are registered to obtain registration information of each conversation person.

[0179] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.

[0180] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0181] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0182] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.

[0183] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0184] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0185] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0186] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0187] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

Claims

1. A voice registration method, characterized in that: The method comprises: Receiving a multi-channel speech signal of a current frame; Extracting sound representations of the multi-channel speech signal of the current frame to determine the sound representation information of each speaker in the current frame; Clustering the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain a representation clustering result of the conversation persons in the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames; Probabilistic analysis is performed based on the representation clustering results to obtain the voiceprint information and spatial position of each conversation person; The voiceprint information and spatial position of each conversation person are registered to obtain registration information of each conversation person.

2. The method according to claim 1, characterized in that The probability analysis is performed based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, including: Performing sound source localization on the multi-channel speech signal of the current frame to obtain a sound source localization result of the multi-channel speech signal of the current frame; Probabilistic analysis is performed in combination with the sound source localization result and the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person.

3. The method according to claim 2, characterized in that The probability analysis is performed in combination with the sound source localization result and the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, including: Determine the voiceprint information of each conversation person according to the cluster center of each conversation person in the characterization clustering result and the probability that the multi-channel speech signal of the current frame belongs to each conversation person; Performing probability statistics calculation on the probability that the multi-channel speech signal of the current frame belongs to each conversation person and the sound source localization result of the multi-channel speech signal of the current frame to obtain the spatial distribution information of each conversation person; The spatial distribution information of each conversation participant is matched with the placeholder information to determine the spatial position of each conversation participant.

4. The method according to claim 3, characterized in that The matching of the spatial distribution information of each conversation person with the placeholder information to determine the spatial position of each conversation person includes: Acquiring the vehicle's place-occupancy information, wherein the place-occupancy information includes various place-occupancy areas; Based on the spatial distribution information of each conversation person, obtain the spatial distribution center position of each conversation person; The spatial distribution center position of each conversation person is matched with the placeholder information to obtain the spatial position of each conversation person.

5. The method according to claim 4, characterized in that The matching of the spatial distribution center position of each conversation person with the placeholder information to obtain the spatial position of each conversation person includes: Matching the spatial distribution center position of each conversation person with the place-occupancy information, and calculating the sum of the Euclidean distances between the matched spatial distribution center position of each conversation person and the center point of each place-occupancy area in the place-occupancy information; When the sum of the Euclidean distances between the spatial distribution center positions of the respective conversation participants and the center points of the respective placeholder areas in the placeholder information is the smallest, the spatial distribution center positions of the respective conversation participants are matched with the placeholder information to obtain the spatial positions of the respective conversation participants.

6. The method according to claim 1, characterized in that The characterization clustering result includes the class center of each conversation person in the current frame and the probability that the multi-channel speech signal of the current frame belongs to each conversation person; The clustering process is performed on the voice characterization information of each conversation person in the current frame and the voice characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain the characterization clustering result of the conversation persons in the current frame, including: Clustering the voice characterization information of each conversation person in the current frame and the voice characterization information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame to obtain the cluster center of each conversation person in the current frame; Based on the sound representation information of each conversation person in the current frame, combined with the class center, class weight and covariance matrix of each conversation person, the probability that the multi-channel speech signal of the current frame belongs to each conversation person is calculated.

7. The method according to claim 1, characterized in that After performing probability analysis based on the characterization clustering result to obtain the voiceprint information and spatial position of each conversation person, the method further includes: Calculating the similarity between the voiceprint information of each conversation person and the voiceprint information registered in the voiceprint database; For the voiceprint information of any conversation person, if the similarity between the voiceprint information of the conversation person and the voiceprint information registered in the voiceprint database is greater than a preset threshold, it is determined that the conversation person and the conversation person corresponding to the voiceprint information registered in the voiceprint database are the same conversation person, and the conversation person is associated with the conversation person corresponding to the voiceprint information registered in the voiceprint database; The voiceprint information and spatial position of each conversation person are registered to obtain registration information of each conversation person, including: The voiceprint information and spatial position of each conversation person are added to the registration information of the conversation person corresponding to the voiceprint information registered in the voiceprint library.

8. The method according to claim 1, characterized in that After registering the voiceprint information and spatial position of each conversation person to obtain the registration information of each conversation person, the method further includes: Real-time monitoring of the voiceprint information and spatial position of each registered conversation person to see if there is any change; If the voiceprint information of each conversation person changes, or the spatial position of each conversation person changes, the registration information of each conversation person is updated.

9. A voice registration device, characterized in that: include: A receiving module, used for receiving a multi-channel speech signal of a current frame; A first determination module is used to extract sound representation information from the multi-channel speech signal of the current frame to determine the sound representation information of each conversation person in the current frame; A first obtaining module is used to perform clustering processing on the sound representation information of each conversation person in the current frame and the sound representation information of each conversation person in the multi-channel speech signal of a preset number of frames before the current frame, to obtain a representation clustering result of the conversation persons in the current frame, wherein the sound representation information of each conversation person in the multi-channel speech signal of the preset number of frames is obtained by performing sound representation extraction on the multi-channel speech signal of the preset number of frames; The second obtaining module is used to perform probability analysis based on the characterization clustering results to obtain the voiceprint information and spatial position of each conversation person; The registration module is used to register the voiceprint information and spatial position of each conversation person to obtain the registration information of each conversation person.

10. An electronic device, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A vehicle, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 8.