Role separation method, electronic device and computer storage medium

By obtaining sound source information and voiceprint features, screening candidate locations and calculating similarity, the problem of character separation errors caused by similar voiceprint features is solved, achieving higher character recognition accuracy.

CN114360574BActive Publication Date: 2025-10-03ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210023782.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-10
Publication Date
2025-10-03
Estimated Expiration
2042-01-10

AI Technical Summary

Technical Problem

During the role separation process, existing technologies have difficulty in effectively distinguishing speakers with similar voiceprint features, resulting in large role separation errors and feedback of erroneous information.

Method used

By obtaining the sound source information and voiceprint features of the target voice data, determining the candidate position based on the sound source information, and calculating the similarity between the voiceprint features of the candidate position and the target voice data, the target role is determined.

Benefits of technology

It improves the accuracy of character separation, reduces the amount of calculation, takes into account the sound source location and voiceprint characteristics, and ensures the accuracy of character recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360574B_ABST
    Figure CN114360574B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for character separation, an electronic device, and a computer storage medium. The method comprises: obtaining sound source information and voiceprint features of target speech data; determining at least one candidate location corresponding to the sound source location based on the sound source information; calculating the similarity between the voiceprint features of the character corresponding to the candidate location and the voiceprint features of the target speech data; and determining a target character corresponding to the target speech data based on the similarity. Embodiments of the present application improve the accuracy of character separation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of speech processing technology, and in particular to a role separation method, an electronic device, and a computer storage medium. Background Art

[0002] In many application scenarios, such as conferences and voice calls, it's necessary to determine the speaker's identity or role based on their voice data in order to provide feedback to the user. Typically, voice data from different roles can be distinguished based on their voiceprint characteristics. However, if two speakers have similar voiceprint characteristics during this role separation process, significant errors can occur during role separation, resulting in incorrect feedback to the user. Summary of the Invention

[0003] In view of this, embodiments of the present application provide a role separation solution to solve some or all of the above problems.

[0004] According to a first aspect of an embodiment of the present application, a role separation method is provided, including: obtaining sound source information and voiceprint features of target voice data; determining at least one candidate position corresponding to the sound source position based on the sound source information; calculating the similarity between the voiceprint features of the role corresponding to the candidate position and the voiceprint features of the target voice data; and determining the target role corresponding to the target voice data based on the similarity.

[0005] According to a second aspect of an embodiment of the present application, a role separation device is provided, including: an acquisition module for acquiring sound source information and voiceprint features of target voice data; a candidate module for determining at least one candidate position corresponding to the sound source position based on the sound source information; a similarity module for calculating the similarity between the voiceprint features of the role corresponding to the candidate position and the voiceprint features of the target voice data; and a role separation module for determining the target role corresponding to the target voice data based on the similarity.

[0006] According to the third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the role separation method described in the first aspect.

[0007] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the role separation method of the first aspect is implemented.

[0008] The character separation solution provided in the embodiments of the present application obtains the sound source information and voiceprint features of the target speech data; determines at least one candidate location corresponding to the sound source location based on the sound source information; calculates the similarity between the voiceprint features of the character corresponding to the candidate location and the voiceprint features of the target speech data; and determines the target character corresponding to the target speech data based on the similarity. Because the candidate locations are first screened based on the sound source locations indicated by the sound source information, the amount of computation is reduced. The similarity between the voiceprint features of the character corresponding to the candidate locations and the voiceprint features of the target speech data is then calculated, and the target character is determined based on the similarity. This takes into account both the sound source location and the voiceprint features, resulting in higher accuracy in character separation. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0010] Figure 1 A schematic diagram of an application scenario of a role separation method provided in Example 1 of the present application;

[0011] Figure 2 A flowchart of a role separation method provided in Example 1 of the present application;

[0012] Figure 3 A flowchart of a role separation method provided in Example 1 of the present application;

[0013] Figure 4 This is a structural diagram of a role separation device provided in Example 2 of the present application;

[0014] Figure 5 This is a structural diagram of an electronic device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0015] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.

[0016] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.

[0017] Example 1

[0018] The first embodiment of the present application provides a role separation method, which is applied to a terminal device. For ease of understanding, the application scenario of the role separation method provided in the first embodiment of the present application is described. Figure 1 As shown, Figure 1 A schematic diagram of a scenario of a role separation method provided in Example 1 of the present application. Figure 1 The illustrated scenario includes an electronic device 101 and a user 102 .

[0019] Figure 1 The scene shown can be a conference room. When the user speaks, the electronic device 101 obtains the sound source information and voiceprint features of the target voice data, determines the candidate position corresponding to the sound source position based on the sound source information, calculates the similarity between the voiceprint features of the character corresponding to the candidate position and the voiceprint features of the target voice data, and determines the role of the speaking user, that is, the target role, based on the similarity.

[0020] The electronic device 101 can access the network, connect to the cloud through the network, and interact with data. In this application, the network includes a local area network (LAN), a wide area network (WAN), and a mobile communication network; such as the World Wide Web (WWW), a long-term evolution (LTE), a 2G network (2nd Generation Mobile Network), a 3G network (3rd Generation Mobile Network), and a 5G network (5th Generation Mobile Network). The cloud can include various devices connected to the network, such as servers, relay devices, and end-to-end (Device-to-Device, D2D) devices. Of course, this is just an exemplary description and does not mean that the application is limited to this.

[0021] Combined with the above Figure 1 In the scenario shown, the first embodiment of the present application provides a role separation method, which is applied to electronic devices. It should be noted that: Figure 1 This is just an example application scenario of the role separation method of this application, and does not mean that the role separation method of this application must be applied to Figure 1 The scene shown, refer to Figure 2 As shown, Figure 2 This is a flowchart of a role separation method provided in Example 1 of the present application, which includes the following steps:

[0022] Step 201: Acquire the sound source information and voiceprint features of the target speech data.

[0023] It should be noted that the target voice data refers to the voice data for which the role is to be determined. This voice data can be divided into at least one data frame based on time. The sound source information indicates the location of the target voice data's sound source, that is, the location of the user emitting the voice. The voiceprint feature indicates the sound wave spectrum characteristics of the user emitting the voice. The user emitting the voice is the user for whom the role is to be determined.

[0024] Alternatively, in one implementation, sound source information can be determined based on sound waves received by a microphone using sound source localization technology. Alternatively, voiceprint features can be extracted from target speech data using a neural network model. This is, of course, merely an example.

[0025] Optionally, when acquiring the initial voice data, the initial voice data can be segmented based on its sound source information, with voice segments from the same sound source location serving as target initial voice data. For example, if the initial voice data contains voice data from two sound source locations, segmentation can be performed at the point where the sound source location changes, resulting in two voice segments. Both of these segments can serve as target voice data for character identification. Each target voice data segment contains only the voice of a single user, further improving the accuracy of character separation.

[0026] Step 202: Determine at least one candidate position corresponding to the sound source position according to the sound source information.

[0027] It should be noted that the at least one candidate location corresponding to the sound source location can be selected based on the sound source information and used to determine whether the character corresponding to the candidate location is the target character's location. For example, in some application scenarios, locations whose azimuth difference from the sound source location is less than or equal to a preset azimuth difference can be used as candidate locations; in other application scenarios, all locations can be used as candidate locations. Of course, this is merely an example.

[0028] Optionally, in one example, it is possible to first determine whether the number of frames of the target voice data is sufficient. If the number of frames of the target voice is too small, the target role can be directly determined based on the azimuth difference between the sound source position and the position of the historical voice data; if the number of frames of the target voice is sufficient, the candidate position can be further determined. For example, determining at least one candidate position corresponding to the sound source position based on the sound source information includes: when the number of frames of the target voice data is greater than the preset number of frames, determining whether the target voice data is the first voice data; if the target voice data is not the first voice data, determining at least one candidate position corresponding to the sound source position based on the sound source information; otherwise, generating a new position as a candidate position based on the sound source information of the target voice data. The preset number of frames can be set according to the specific situation. Optionally, the preset number of frames can be greater than or equal to 50, or the preset number of frames can be greater than or equal to 100, etc.

[0029] Optionally, based on the above example, in one implementation, if the target voice data is not the first voice data, at least one candidate position corresponding to the sound source position is determined based on the sound source information, including: if the target voice data is not the first voice data, calculating the azimuth change difference of the target voice data relative to the position closest in azimuth based on the sound source information; if the azimuth change difference is greater than the preset change difference, then the existing position is determined as the candidate position; otherwise, the position closest in azimuth is determined as the candidate position. If the azimuth change difference is greater than the preset change difference, it means that the position closest in azimuth is far away from the sound source position in space and is not the same user. In this case, it is very likely that the user corresponding to the target voice data has moved. Therefore, other existing positions are further screened as candidate positions to ensure that the accuracy of role determination can still be high when the user position moves.

[0030] In one example, if the target voice data is not the first voice data, and the orientation change difference is greater than the preset change difference, then as described above, the existing position will be determined as a candidate position, and then the following steps 203 and 204 will be performed to determine the target role based on the similarity between the voiceprint features of the role corresponding to the candidate position and the voiceprint features of the target voice data. As described above, there may be a situation where the user moves the position. In this case, the correspondence between the target voice data and the candidate position can be recorded. In this way, for a voice segment that includes multiple roles and has multiple target voice data, after determining the target role corresponding to each target voice data, it is possible to further determine which target voice data a specific target role corresponds to in the voice segment and how the positions of these target voice data have changed based on the correspondence between the target role and the candidate position. That is, after determining the target role, the correspondence between the target role and the candidate position of the voiceprint feature with the highest similarity can be recorded; based on the correspondence, it is determined whether the candidate position in multiple (two or more) target voice data corresponding to the target role (including the current target voice data and historical target voice data corresponding to the target role) has changed; if a change has occurred, the position change information of the target role can be determined based on the change.

[0031] Step 203: Calculate the similarity between the voiceprint feature of the character corresponding to the candidate position and the voiceprint feature of the target voice data.

[0032] It should be noted that the similarity of voiceprint features can be obtained by calculating the Euclidean distance between two voiceprint features, or by scoring through Probabilistic Linear Discriminant Analysis (PLDA).

[0033] Step 204: Determine the target character corresponding to the target voice data based on the similarity.

[0034] It should be noted that the higher the similarity, the greater the likelihood that the character corresponding to the candidate position is the same character as the character corresponding to the target voice data. Therefore, the target character can be determined based on the similarity. For example, determining the target character corresponding to the target voice data based on the similarity includes: determining the character corresponding to the candidate position with the most similar voiceprint features as the target character. Determining the character with the most similar voiceprint features to the target voice data as the target character more accurately separates the target character.

[0035] Based on the example in step 202 above, two scenarios are listed here to illustrate how to determine the target role.

[0036] Optionally, in the first scenario, the target character corresponding to the target voice data is determined based on similarity, including: if the target voice data is not the first voice data, calculating the azimuth change difference of the target voice data relative to the position closest in azimuth based on the sound source information; if the azimuth change difference is less than or equal to a preset change difference, and the similarity is greater than the preset similarity, determining the character corresponding to the similarity as the target character; if the azimuth change difference is less than or equal to the preset change difference, and the similarity is less than or equal to the preset similarity, calculating the similarity between the voiceprint features corresponding to other locations within the candidate location's area and the voiceprint features of the target voice data, and determining the character corresponding to the voiceprint features with a similarity greater than the preset similarity as the target character. If the target voice data is not the first voice data, it indicates that historical voice data already exists, that is, other characters already exist. Therefore, it is necessary to determine whether the character corresponding to the target voice data is another character that has already spoken to avoid omissions. If the orientation change difference is less than or equal to the preset change difference, it means that the position closest to the sound source position is very close to the sound source position, and it is very likely that it is the same character. However, if the orientation change difference is greater than the preset change difference, it means that the position closest to the sound source position is far away from the sound source position, and it is very likely that the speaker has moved, and it is necessary to screen other positions in the area where the candidate position is located. The orientation change difference can be represented by the angle formed by the line segment from the sound source position to the reference point and the line segment from the position closest to the sound source position (candidate position) to the reference point. For example, the preset change difference can be 40 degrees.

[0037] Whether the speaker has moved can be determined based on the correspondence between the target voice data and the candidate positions. In this case, each time the target character is determined, the correspondence between the target voice data and the closest position should be recorded. Whether the target character has moved can be determined based on whether the position of the target character has changed in different target voice data.

[0038] For example, a speech segment containing multiple characters can be segmented based on the changes in the sound source position as described above, resulting in multiple speech segments. In this example, these segments are set to include Segment 1, Segment 2, and Segment 3, each of which can be used as a target speech data. Alternatively, the speech segment can be segmented based on changes in voiceprint features. For example, this segmentation is also set to be segmented into Segment 1, Segment 2, and Segment 3.

[0039] Assume that, through the above process, the target character of voice segment 1 is determined to be A, and its corresponding closest position X is recorded; the target character of voice segment 2 is determined to be B, and its corresponding closest position Y is recorded; the target character of voice segment 3 is determined to be A, and its corresponding closest position Z is recorded. It can be seen that in this voice segment, target character A speaks twice and moves.

[0040] In the first scenario, further optionally, the method further includes: if the similarity of the voiceprint features of other locations within the candidate location is less than or equal to a preset similarity, then calculating the similarity of the voiceprint features corresponding to the locations in the other locations with the voiceprint features of the target voice data, and determining the character corresponding to the voiceprint features with a similarity greater than the preset similarity as the target character; if the similarity of the voiceprint features corresponding to the locations in the other locations is less than or equal to the preset similarity, then generating a new character for the target voice data as the target character. In the first scenario, the location closest in direction (i.e., the candidate location) is first determined; if the direction change difference of the closest location is greater than the preset change difference, then the range is expanded to determine other locations in the area where the closest location is located; if the similarity of the voiceprint features corresponding to other locations in the area is less than or equal to the preset similarity, then the range is further expanded to determine the locations of other areas until the target character is determined. In this way, the range is expanded layer by layer based on the sound source location, which ensures accuracy and avoids omissions. It should also be noted that the area can be sector-shaped and can be distinguished at different angles. For example, a sector corresponding to 45 degrees is one area, and a scene can be divided into 8 areas. There may be at least one location in an area, or there may be no set location, and new locations can be gradually established based on user comments.

[0041] Based on this, a feasible role separation scheme of the embodiment of the present application can be implemented as follows: obtaining the sound source information and voiceprint features of the target voice data; determining the spatial partition to which the sound source position indicated by the sound source information belongs, and determining at least one candidate position corresponding to the sound source position in the spatial partition; wherein the spatial partition is one of the multiple spatial regions formed after the physical space where the speaker corresponding to the target voice data is located is spatially divided according to a preset angle; calculating the similarity between the voiceprint features of the role corresponding to the candidate position and the voiceprint features of the target voice data; and determining the target role corresponding to the target voice data based on the similarity. The preset angle can be set by those skilled in the art according to actual needs, and the embodiment of the present application does not impose any restrictions on this.

[0042] Further optionally, the determination of at least one candidate position corresponding to the sound source position in the spatial partition can be implemented as follows: judging whether there is a candidate position corresponding to the sound source position in the spatial partition; if so, determining the candidate position as the candidate position corresponding to the sound source position in the spatial partition; if not, establishing a candidate position in the spatial partition based on the sound source position.

[0043] Refer again Figure 1 , Figure 1 In the example, the speaker's physical space is divided into three spatial regions at 45-degree angles, i.e., eight spatial partitions. Assume that the spatial partition to which the corresponding sound source position belongs is determined to be the first partition according to the sound source information of the target speech data, i.e. Figure 1 In the partition where the "+" circle is located, when determining the candidate position, first determine the candidate position corresponding to the sound source position from the first partition (there may be one or more), Figure 1 If there is a candidate position in the first partition that is the same as the sound source position, the similarity between the voiceprint features of the character corresponding to the candidate position and the voiceprint features of the target voice data can be calculated first; then the target character corresponding to the target voice data is determined based on the similarity. Of course, if the similarities corresponding to the candidate positions in the same spatial partition are all low, the similarities between the voiceprint features of the characters corresponding to the candidate positions in other spatial partitions and the voiceprint features of the target voice data can be calculated, such as Figure 1 The candidate positions in the lower partition adjacent to the first partition are shown in .

[0044] Assuming that there is no candidate position in the first partition, in this case, a new candidate position can be created in the first partition based on the sound source position. For example, the sound source position can be directly created as a candidate position for subsequent use when needed.

[0045] Through the above method, the target role can be determined more accurately and effectively, and the candidate positions can be supplemented and improved to improve the overall efficiency of the plan.

[0046] Optionally, in the second scenario, the method further includes: when the target voice data has a frame count of less than or equal to a preset frame count, determining candidate voice data in the historical voice data that is closest in orientation to the target voice data based on the sound source information; calculating the orientation difference between the target voice data and the candidate voice data, and if the orientation difference is less than a preset threshold, determining the character corresponding to the candidate voice data as the target character. If the target voice data has a frame count of less than or equal to the preset frame count, it may not be possible to determine the character based on the similarity of the voiceprint features because the small number of frames makes the calculated similarity less accurate. Therefore, determination can be made directly based on the orientation of the historical voice data. It should be noted that, in this application, the orientation difference between the target voice data and the candidate voice data refers to the sound source position corresponding to the target voice data. The orientation difference between the position corresponding to the candidate voice data can also be understood as the orientation change difference. The orientation difference can be represented by the angle formed by a line segment from the sound source position to a reference point and a line segment from the position corresponding to the candidate voice data closest in orientation to the reference point. For example, the preset threshold can be 5 degrees.

[0047] In combination with the role separation method described in steps 201-204 above, a specific application scenario is listed here for detailed description. Figure 3 As shown, Figure 3 This is a flowchart of a role separation method provided in Example 1 of the present application. After obtaining the target voice data, it is first determined whether the frame number of the target voice data is greater than a preset frame number (the preset frame number can be 100); if it is less than or equal to the preset frame number, it is compared with the historical voice data to determine the candidate voice data with the closest orientation, and it is determined whether the orientation difference between the candidate voice data and the target voice data is less than a preset threshold (the preset threshold is 5). If it is less than the preset threshold, the role corresponding to the candidate voice data is the target role. If it is greater than or equal to the preset threshold, no determination can be made.

[0048] If the number of frames of the target voice data is greater than the preset number of frames, it is further determined whether the target voice data is the first voice data. If it is the first voice data, a new character is established for the target voice data as the target character. A new position and a new area can also be established based on the sound source position of the target voice data. If the target voice data is not the first voice data, the positions of all areas are traversed to determine the position closest to the sound source position of the target voice data, and the azimuth change difference between the sound source position and the position closest to its azimuth is calculated. It is determined whether the azimuth change difference is greater than the preset change difference (the azimuth change difference can be 40 degrees). If the azimuth change difference is greater than the preset change difference, the voiceprint features of the target voice data are compared with the voiceprint features of all positions in other areas to calculate the similarity. If the similarity is greater than the preset similarity, the character at the position corresponding to the similarity is determined as the target character. If the similarity is less than or equal to the preset similarity, a new character is generated for the target voice data as the target character. A new position and a new area can also be generated for the target voice data.

[0049] If the direction change difference is less than or equal to the preset change difference, it can be further determined whether the direction change difference is less than the difference lower limit (the difference lower limit can be 10 degrees). If it is less than the difference lower limit, it can be determined that the character corresponding to the position with the closest direction is the target character. If the direction change difference is greater than or equal to the difference lower limit, the voiceprint features corresponding to the position with the closest direction are compared with the voiceprint features of the target voice data, and the similarity is calculated to determine whether the similarity is greater than the preset similarity. If it is, the character corresponding to the position with the closest direction is the target character. If the similarity is less than or equal to the preset similarity, other positions in the area where the position with the closest direction is located are taken as candidate positions to expand the comparison range.

[0050] Calculate the similarity between the voiceprint features corresponding to the candidate position and the voiceprint features of the target voice data. If the similarity is greater than the preset similarity, the character at the candidate position corresponding to the similarity is taken as the target character. If the similarity is less than or equal to the preset similarity, the positions of all other areas are taken as candidate positions to further expand the comparison range. Calculate the similarity between the voiceprint features corresponding to the candidate position and the voiceprint features of the target voice data. If the similarity is greater than the preset similarity, the character at the candidate position corresponding to the similarity is taken as the target character. If all candidate positions in the area have been compared and there is no position with a similarity greater than the preset similarity, a new character is generated for the target voice data as the target character, and the new position is set based on the sound source position. It should also be noted that if there are voiceprint features of more than two candidate positions in an area, and their similarity with the voiceprint features of the target voice data is greater than the preset similarity, then among these candidate positions, the character corresponding to the candidate position with the greatest similarity is determined as the target character.

[0051] The character separation method provided in the embodiments of the present application obtains the sound source information and voiceprint features of the target speech data; determines at least one candidate location corresponding to the sound source location based on the sound source information; calculates the similarity between the voiceprint features of the character corresponding to the candidate location and the voiceprint features of the target speech data; and determines the target character corresponding to the target speech data based on the similarity. Because the candidate locations are first screened based on the sound source locations indicated by the sound source information, the amount of computation is reduced. The similarity between the voiceprint features of the character corresponding to the candidate locations and the voiceprint features of the target speech data is then calculated, and the target character is determined based on the similarity. This method takes both the sound source location and the voiceprint features into consideration, resulting in higher accuracy in character separation.

[0052] Example 2

[0053] Based on the method described in the above embodiment 1, the second embodiment of the present application provides a role separation device for executing the method described in the above embodiment 1, referring to Figure 4 As shown, the role separation device 40 includes:

[0054] Acquisition module 401, used to obtain sound source information and voiceprint features of target speech data;

[0055] A candidate module 402 is configured to determine at least one candidate position corresponding to the sound source position based on the sound source information;

[0056] Similarity module 403, used to calculate the similarity between the voiceprint feature of the character corresponding to the candidate position and the voiceprint feature of the target voice data;

[0057] The role separation module 404 is configured to determine the target role corresponding to the target speech data according to the similarity.

[0058] Optionally, in one embodiment, the role separation module 404 is configured to determine, among the roles corresponding to the candidate positions, the role with the greatest similarity in voiceprint features as the target role.

[0059] Optionally, in one embodiment, the candidate module 402 is used to determine whether the target voice data is the first voice data when the number of frames of the target voice data is greater than a preset number of frames; if the target voice data is not the first voice data, determine at least one candidate position corresponding to the sound source position based on the sound source information; otherwise, generate a new position as a candidate position based on the sound source information of the target voice data.

[0060] Optionally, in one embodiment, the candidate module 402 is used to calculate the azimuth change difference of the target voice data relative to the position closest in azimuth based on the sound source information if the target voice data is not the first voice data; if the azimuth change difference is greater than the preset change difference, the existing position is determined as the candidate position; otherwise, the position closest in azimuth is determined as the candidate position.

[0061] Optionally, in one embodiment, the role separation module 404 is used to calculate the azimuth change difference of the target voice data relative to the position closest to the azimuth based on the sound source information if the target voice data is not the first voice data; if the azimuth change difference is less than or equal to the preset change difference, and the similarity is greater than the preset similarity, determine the role corresponding to the similarity as the target role; if the azimuth change difference is less than or equal to the preset change difference, and the similarity is less than or equal to the preset similarity, calculate the similarity between the voiceprint features corresponding to other positions in the area where the candidate position is located and the voiceprint features of the target voice data, and determine the role corresponding to the voiceprint features with a similarity greater than the preset similarity as the target role.

[0062] Optionally, in one embodiment, the role separation module 404 is further used to calculate the similarity between the voiceprint features corresponding to the locations in the area where the candidate location is located and the voiceprint features of the target voice data if the similarity of the voiceprint features is less than or equal to a preset similarity, and determine the role corresponding to the voiceprint features with a similarity greater than the preset similarity as the target role; if the similarity of the voiceprint features corresponding to the locations in the other areas is less than or equal to the preset similarity, generate a new role for the target voice data as the target role.

[0063] Optionally, in one embodiment, the role separation module 404 is also used to determine, when the number of frames of the target voice data is less than or equal to a preset number of frames, candidate voice data that is closest in orientation to the target voice data in the historical voice data based on the sound source information; calculate the orientation difference between the target voice data and the candidate voice data, and if the orientation difference is less than a preset threshold, determine the role corresponding to the candidate voice data as the target role.

[0064] The apparatus provided in an embodiment of the present application obtains the sound source information and voiceprint features of target speech data; determines at least one candidate location corresponding to the sound source location based on the sound source information; calculates the similarity between the voiceprint features of the character corresponding to the candidate location and the voiceprint features of the target speech data; and determines the target character corresponding to the target speech data based on the similarity. Because the candidate locations are first screened based on the sound source locations indicated by the sound source information, the amount of computation is reduced. The similarity between the voiceprint features of the character corresponding to the candidate locations and the voiceprint features of the target speech data is then calculated, and the target character is determined based on the similarity. This takes both the sound source location and the voiceprint features into account, resulting in higher accuracy in character separation.

[0065] Example 3

[0066] Based on the method described in the above embodiment 1, the third embodiment of the present application provides an electronic device for executing any of the methods described in the above embodiment 1, referring to Figure 5 As shown, Figure 5 This is a structural diagram of an electronic device provided in Example 3 of the present application. The specific embodiments of the present application do not limit the specific implementation of the electronic device.

[0067] like Figure 5 As shown, the electronic device may include: a processor (processor) 502 , a communication interface (Communications Interface) 504 , a memory (memory) 506 , and a communication bus 508 .

[0068] in:

[0069] The processor 502 , the communication interface 504 , and the memory 506 communicate with each other via a communication bus 508 .

[0070] The communication interface 504 is used to communicate with other electronic devices such as terminal devices or servers.

[0071] The processor 502 is configured to execute the program 510 , and specifically may execute the relevant steps in the above method embodiment.

[0072] Specifically, the program 510 may include program codes, which include computer operation instructions.

[0073] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the electronic device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0074] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0075] The program 510 may be specifically configured to enable the processor 502 to execute any of the methods in the aforementioned embodiments.

[0076] The specific implementation of each step in program 510 can be found in the corresponding descriptions of the corresponding steps and units in the above-mentioned speed detection method embodiment, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the above-mentioned method embodiment, and will not be repeated here.

[0077] The electronic device provided in an embodiment of the present application obtains the sound source information and voiceprint features of target voice data; determines at least one candidate location corresponding to the sound source location based on the sound source information; calculates the similarity between the voiceprint features of the character corresponding to the candidate location and the voiceprint features of the target voice data; and determines the target character corresponding to the target voice data based on the similarity. Because the candidate locations are first screened based on the sound source locations indicated by the sound source information, the amount of computation is reduced. The similarity between the voiceprint features of the character corresponding to the candidate locations and the voiceprint features of the target voice data is then calculated, and the target character is determined based on the similarity. This takes both the sound source location and the voiceprint features into account, resulting in higher accuracy in character separation.

[0078] Example 4

[0079] Based on the method described in the above embodiment 1, embodiment 4 of the present application provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, any of the methods described in embodiment 1 is implemented.

[0080] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0081] The above-described method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium downloaded via a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor or hardware, the role separation method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the role separation method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the role separation method shown herein.

[0082] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0083] The above implementation methods are only used to illustrate the embodiments of the present application, and are not intended to limit the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims.

Claims

1. A role separation method, comprising: Obtaining sound source information and voiceprint features of target voice data; determining at least one candidate position corresponding to the sound source position according to the sound source information, wherein the candidate position is used to represent a position screened according to the sound source information; Calculating the similarity between the voiceprint feature of the character corresponding to the candidate position and the voiceprint feature of the target voice data; Determining a target role corresponding to the target voice data based on the similarity; Among them, determining at least one candidate position corresponding to the sound source position based on the sound source information includes: when the number of frames of the target voice data is greater than the preset number of frames, determining whether the target voice data is the first voice data; if the target voice data is not the first voice data, determining at least one candidate position corresponding to the sound source position based on the sound source information.

2. The method according to claim 1, wherein Determining a target role corresponding to the target voice data according to the similarity includes: Among the characters corresponding to the candidate positions, the character with the greatest similarity in voiceprint features is determined as the target character.

3. The method according to claim 1, wherein The method further comprises: If the target voice data is the first voice data, a new position is generated as a candidate position according to the sound source information of the target voice data.

4. The method according to claim 3, wherein: If the target voice data is not the first voice data, determining at least one candidate position corresponding to the sound source position according to the sound source information includes: If the target voice data is not the first voice data, calculating the direction change difference of the target voice data relative to the position closest to the direction according to the sound source information; If the orientation change difference is greater than a preset change difference, the existing position is determined as the candidate position; otherwise, the position with the closest orientation is determined as the candidate position.

5. The method according to claim 3, wherein: The determining, based on the similarity, a target role corresponding to the target voice data includes: If the target voice data is not the first voice data, calculating the direction change difference of the target voice data relative to the position closest to the direction according to the sound source information; If the orientation change difference is less than or equal to a preset change difference, and the similarity is greater than a preset similarity, determining the character corresponding to the similarity as the target character; If the orientation change difference is less than or equal to the preset change difference, and the similarity is less than or equal to the preset similarity, then the similarity between the voiceprint features corresponding to other positions in the area where the candidate position is located and the voiceprint features of the target voice data is calculated, and the role corresponding to the voiceprint features with a similarity greater than the preset similarity is determined as the target role.

6. The method according to claim 5, wherein: The method further comprises: If the similarities of the voiceprint features of other locations within the area where the candidate location is located are all less than or equal to the preset similarity, then the similarities of the voiceprint features corresponding to the locations in the other areas with the voiceprint features of the target voice data are calculated, and the character corresponding to the voiceprint feature having a similarity greater than the preset similarity is determined as the target character; If the similarities of the voiceprint features corresponding to the positions in other areas are all less than or equal to the preset similarity, a new role is generated for the target voice data as the target role.

7. The method according to claim 3, wherein: The method further comprises: When the number of frames of the target voice data is less than or equal to the preset number of frames, determining candidate voice data closest in position to the target voice data based on the historical voice data of the sound source information; The orientation difference between the target voice data and the candidate voice data is calculated, and if the orientation difference is less than a preset threshold, the role corresponding to the candidate voice data is determined as the target role.

8. The method according to claim 1, wherein The method further comprises: Recording the correspondence between the target character and the candidate position of the voiceprint feature with the highest similarity; According to the corresponding relationship, determining whether the candidate positions in the plurality of target voice data corresponding to the target character have changed; If a change occurs, the position change information of the target character is determined according to the change.

9. A role separation method, comprising: Obtaining sound source information and voiceprint features of target voice data; Determining a spatial partition to which the sound source position indicated by the sound source information belongs, and determining at least one candidate position corresponding to the sound source position in the spatial partition; wherein the spatial partition is one of a plurality of spatial regions formed by spatially dividing the speaker's physical space corresponding to the target speech data according to a preset angle, and the candidate position is used to represent a position selected based on the sound source information; Calculating the similarity between the voiceprint feature of the character corresponding to the candidate position and the voiceprint feature of the target voice data; Determining a target role corresponding to the target voice data based on the similarity; Among them, determining the spatial partition to which the sound source position indicated by the sound source information belongs, and determining at least one candidate position corresponding to the sound source position in the spatial partition, includes: when the number of frames of the target voice data is greater than the preset number of frames, determining whether the target voice data is the first voice data; if the target voice data is not the first voice data, determining at least one candidate position corresponding to the sound source position in the spatial partition according to the spatial partition to which the sound source position indicated by the sound source information belongs.

10. The method according to claim 9, wherein: The determining of at least one candidate position corresponding to the sound source position in the spatial partition includes: Determining whether there is a candidate position corresponding to the sound source position in the spatial partition; If yes, determining the candidate position as the candidate position corresponding to the sound source position in the spatial partition; If not, a candidate position is established in the spatial partition according to the sound source position.

11. An electronic device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the role separation method according to any one of claims 1 to 10.

12. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the role separation method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Object recognition method and device, storage medium and terminal

    CN108305615A

  • Sound source tracking method and device based on biological characteristics, equipment and storage medium

    CN109754811A