Voice data processing method and device

Through a voice data processing method, voice data is automatically recognized and processed to obtain the voiceprint data of the target user, solving the problem of unrecognized during the voiceprint recognition process, and realizing unattended voiceprint registration and update.

CN120071940APending Publication Date: 2025-05-30LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237569.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

During the voiceprint recognition process, due to the change of the voiceprint moment, the problem of unrecognition is caused by the user's manual operation.

Method used

A voice data processing method is provided, which determines the target user in response to the acquisition of voice data, and identifies the matching relationship between the voice data and the target user. According to the recognition results, single or multiple voice data processing is performed, the voiceprint data of the target user is obtained, and voiceprint registration is performed.

Benefits of technology

It realizes automatic voiceprint registration and update without user participation, simplifies operations and improves the automation of voiceprint recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071940A_ABST
    Figure CN120071940A_ABST
Patent Text Reader

Abstract

The invention provides a voice data processing method and device, and the method comprises the steps: responding to the obtaining of voice data at a first moment, determining a target user, and recognizing the matching relation between the voice data and the target user; in response to the fact that the recognition result is the first recognition result, determining that the voice data is single-person voice data or multi-person voice data; when the voice data is the single-person voice data, performing first processing on the voice data; when the voice data is multi-user voice data, second processing is carried out on the voice data, the second processing is different from the first processing, and the first processing and the second processing are both used for obtaining voiceprint data corresponding to the target user; and performing voiceprint registration of the target user based on the first processing result or the second processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and particularly to a method and device for processing voice data. Background Art

[0002] Voiceprint refers to the characteristics and patterns of a voice signal in terms of frequency, amplitude, and time domain, and is unique like a fingerprint. It is determined by the structures and usage habits of vocal organs such as the vocal cords, oral cavity, and nasal cavity. Each person's pronunciation and speech are the result of the combined operation of multiple organs such as the nasal cavity, tongue, oral cavity, vocal tract, chest, and lungs. There are slight differences in the frequency, timbre, intonation, and even accent of different people's speech, and these differences ultimately form completely different voiceprint maps.

[0003] Due to the uniqueness of voiceprints, many devices and programs use voiceprints for user identification during user recognition. However, affected by factors such as a person's mood or physical condition, voiceprints actually change all the time. Therefore, the phenomenon of unrecognizability often occurs during voiceprint recognition, and in such cases, manual operations by the user are required. Summary of the Invention

[0004] Embodiments of the present application provide a method for processing voice data, including:

[0005] In response to obtaining the voice data at a first moment, determining a target user, and identifying the matching relationship between the voice data and the target user;

[0006] In response to the recognition result being a first recognition result, determining whether the voice data is single-person voice data or multi-person voice data;

[0007] When it is single-person voice data, performing a first process on the voice data;

[0008] When it is multi-person voice data, performing a second process on the voice data, where the second process is different from the first process, and both the first process and the second process are used to obtain voiceprint data corresponding to the target user;

[0009] Performing voiceprint registration of the target user based on the first processing result or the second processing result.

[0010] In one embodiment, the method further includes:

[0011] In response to obtaining the voice data at a second moment, performing a first process or a second process on the voice data to obtain a first processing result or a second processing result corresponding to the second moment;

[0012] Performing corresponding matching on the first processing results and the second processing results at the second moment and the first moment to obtain a first matching result;

[0013] If the similarity represented by the first matching result meets the requirements, voiceprint registration of the target user is performed based on the first processing result or the second processing result at the second moment.

[0014] In one embodiment, the method further includes:

[0015] If the similarity represented by the first matching result does not meet the requirements, obtain the voice data at the third moment, and perform the first processing or the second processing on the voice data at the third moment to obtain the first processing result or the second processing result corresponding to the third moment;

[0016] Perform corresponding matching on the first processing results and the second processing results at the first moment, the second moment, and the third moment to obtain different second matching results;

[0017] Perform voiceprint registration of the target user based on the first processing result or the second processing result at two moments corresponding to the second matching result with the similarity meeting the requirements.

[0018] In one embodiment, the determining the target user includes:

[0019] Obtain device information, and determine the target user based on the device information; or

[0020] Obtain the registration information of the target program, and determine the target user based on the registration information; or

[0021] Determine the target user based on the input identity information; or

[0022] Determine the target user based on the stored historical voiceprint registration data.

[0023] In one embodiment, the responding to the obtaining of the voice data at the first moment includes:

[0024] Responding to the obtaining of the voice data within the specified sound frequency range at the first moment; or

[0025] Responding to the obtaining of the voice data with the volume meeting the requirements at the first moment; or

[0026] When it is determined that there is audio input and output, responding to the obtaining of the voice data at the first moment; or

[0027] When it is determined that the target function is started, responding to the obtaining of the voice data at the first moment; or

[0028] When it is determined that the target program is started, responding to the obtaining of the voice data at the first moment.

[0029] In one embodiment, the performing the first processing on the voice data includes:

[0030] Determine the volume and duration of the voice data;

[0031] When the volume and duration of the voice data meet the first requirement, determine the voiceprint data corresponding to the target user based on the voice data.

[0032] In one embodiment, perform a second process on the voice data, including:

[0033] Perform multi-person audio separation calculation on the voice data, and extract the target voice data corresponding to the target user;

[0034] Determine the volume and duration of the voice data;

[0035] When the volume and duration of the voice data meet the first requirement, determine the voiceprint data corresponding to the target user based on the voice data.

[0036] In one embodiment, the performing multi-person audio separation calculation on the voice data and extracting the target voice data corresponding to the target user includes:

[0037] Determine the collection angles corresponding to the multiple collected voice data respectively;

[0038] Extract the voice data corresponding to the target angle as the target voice data of the target user; or

[0039] Determine the collection distances corresponding to the multiple collected voice data respectively;

[0040] Extract the voice data corresponding to the target distance as the target voice data of the target user

[0041] Determine the audio parameters corresponding to the multiple collected voice data respectively;

[0042] Extract the voice data with target audio parameters as the target voice data of the target user.

[0043] In one embodiment, the performing voiceprint registration of the target user based on the first processing result or the second processing result includes:

[0044] In response to an input instruction, perform voiceprint registration of the target user based on the first processing result or the second processing result; or

[0045] In response to an input instruction, update the historical voiceprint registration data of the target user based on the first processing result or the second processing result.

[0046] Another embodiment of the present application also provides a voice data processing device, including:

[0047] The first response module is used to determine the target user in response to the acquisition of voice data at the first moment, and identify the matching relationship between the voice data and the target user;

[0048] The second response module is used to determine whether the voice data is single-person voice data or multi-person voice data in response to the recognition result being the first recognition result;

[0049] The first processing module is used to perform a first processing on the voice data when it is single-person voice data;

[0050] The second processing module is used to perform a second processing on the voice data when it is multi-person voice data. The second processing is different from the first processing, and both the first processing and the second processing are used to obtain the voiceprint data corresponding to the target user;

[0051] The registration module is used to perform voiceprint registration of the target user based on the first processing result or the second processing result.

[0052] Other features and advantages of the present application will be described in the subsequent description, and, in part, will be obvious from the description, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings.

[0053] The following will further describe the technical solutions of the present application in detail through the drawings and embodiments. Description of the Drawings

[0054] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 It is a schematic flowchart of the voice data processing method in the embodiment of the present application.

[0056] Figure 2 It is a schematic flowchart of the voice data processing method in another embodiment of the present application.

[0057] Figure 3 It is a schematic application flowchart of the voice data processing method in the embodiment of the present application.

[0058] Figure 4 It is a structural block diagram of the voice data processing device in the embodiment of the present application. Detailed Embodiments

[0059] Next, specific embodiments of the present application will be described in detail with reference to the accompanying drawings, but this is not a limitation of the present application.

[0060] It should be understood that various modifications can be made to the embodiments disclosed herein. Therefore, the following description should not be regarded as restrictive, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope of the present disclosure.

[0061] The accompanying drawings, which are included in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the general description of the present disclosure given above and the detailed description of the embodiments given below, serve to explain the principles of the present disclosure.

[0062] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non - limiting examples with reference to the accompanying drawings.

[0063] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application, which have the features as described in the claims and thus are all within the scope of protection defined thereby.

[0064] When combined with the accompanying drawings, the above - mentioned and other aspects, features, and advantages of the present disclosure will become more apparent in view of the following detailed description.

[0065] Hereinafter, specific embodiments of the present disclosure will be described with reference to the accompanying drawings; however, it should be understood that the disclosed embodiments are merely examples of the present disclosure and can be implemented in various ways. Well - known and / or repetitive functions and structures are not described in detail to avoid obscuring the present disclosure with unnecessary or redundant details. Therefore, the specific structural and functional details disclosed herein are not intended to be limiting, but merely serve as a basis for the claims and a representative basis for teaching those skilled in the art to use the present disclosure in substantially any suitable detailed structure in a variety of ways.

[0066] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present disclosure.

[0067] Next, embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0068] As Figure 1 shown, an embodiment of the present invention provides a method for processing voice data, including:

[0069] S1: In response to obtaining the voice data at the first moment, determine the target user, and identify the matching relationship between the voice data and the target user;

[0070] S2: In response to the recognition result being the first recognition result, determine whether the voice data is single-person voice data or multi-person voice data;

[0071] S3: When it is single-person voice data, perform a first processing on the voice data;

[0072] S4: When it is multi-person voice data, perform a second processing on the voice data, where the second processing is different from the first processing, and both the first processing and the second processing are used to obtain the voiceprint data corresponding to the target user;

[0073] S5: Perform voiceprint registration of the target user based on the first processing result or the second processing result.

[0074] In this embodiment, the method (program) is automatically started and run when it is determined that the voice data is obtained at the first moment, the target user is determined, and whether there is a matching relationship between the target user and the voice data is identified. For example, after determining the target user, the voiceprint data, voice data, etc. stored locally are retrieved, and the retrieved data is compared with the obtained voice data to determine whether there is a matching relationship between the voice data and the target user, and the corresponding recognition result is obtained. When the recognition result is the first recognition result, it indicates that voiceprint registration of the target user is required. The voiceprint registration can be the first voiceprint registration or the registration data from the past can be updated for re-voiceprint registration, and the specific situation is uncertain. In response to the first recognition result, the system will determine whether the voice data is single-person voice data or multi-person voice data. When it is single-person voice data, the voice data is subjected to the first processing. When it is multi-person voice data, the voice data is subjected to the second processing. The process of the second processing is different from that of the first processing, but both the first processing and the second processing are processes for obtaining the voiceprint data corresponding to the target user. After completing the first processing or the second processing, the first processing result or the second processing result is obtained. At this time, both the first processing result and the second processing result contain the voiceprint data corresponding to the target user. Therefore, the program can complete the voiceprint registration of the target user based on the first processing result or the second processing result.

[0075] Through the method of the above embodiments, the program can be automatically started and run when voice data is obtained, determine whether the current target user has the need for voiceprint registration. If it is determined that there is such a need, the voice data collected is automatically processed for the target user to obtain a processing result containing the voiceprint data of the target user. Then, based on this processing result, the system can automatically complete the first voiceprint registration for the user or update the registered voiceprint data. The above process does not require user participation and is automatically completed by the system and the program, greatly simplifying the user operation and realizing automatic update and registration of the user's voiceprint for subsequent automatic user identification.

[0076] In one embodiment, the obtaining of the voice data in response to the first moment includes:

[0077] Obtaining the voice data within a specified sound frequency range at the first moment; or

[0078] Obtaining the voice data with the volume meeting the requirements at the first moment; or

[0079] When it is determined that there is audio input and output, obtaining the voice data in response to the first moment; or

[0080] When it is determined that the target function is started, obtaining the voice data in response to the first moment; or

[0081] When it is determined that the target program is started, obtaining the voice data in response to the first moment.

[0082] For example, a sound frequency range is preset, which can but is not limited to be determined according to the audio range during the user's normal communication, or can also be directly determined based on the audio range that the human body can output and receive (BF audio range). When the audio of the voice data collected at the first moment is within this sound frequency range, an automatic response is made. Or

[0083] Determine the volume of the obtained voice data. If the volume meets the preset volume requirements, an automatic response is made. Or

[0084] After obtaining the voice data, determine the current scene of the system. If the scene is that the speaker has output audio data and the microphone has received voice data, that is, there is audio input and output, then determine that the scene meets the requirements, so an automatic response is made to the obtaining of the voice data. Or

[0085] When the voice data is obtained, determine whether the system has started the target function. The target function can but is not limited to be a call function, a recording function, a voiceprint recognition function, etc. If so, an automatic response is made. Or determine whether the system has started the target function. If so, monitor whether the voice data is obtained. If so, an automatic response is made. Or

[0086] When voice data is obtained, determine whether the system has launched a target program, which may but is not limited to a conference program, a communication program, etc. If so, make an automatic response. Or determine whether the system has launched a target program. If so, monitor whether voice data has been obtained. If so, make an automatic response.

[0087] Further, after making a response to the voice data, the program starts, and further determines the target user to determine whether there is a matching relationship between the obtained voice data and the target user. The target user refers to the user who is currently using the device and inputs voice data. The program needs to determine the identity of this target user. The determination process includes:

[0088] Obtain device information and determine the target user based on the device information; or

[0089] Obtain the registration information of the target program and determine the target user based on the registration information; or

[0090] Determine the target user based on the input identity information; or

[0091] Determine the target user based on the stored historical voiceprint registration data.

[0092] Exemplarily, the program can default that the user currently using the electronic device is the device owner. Therefore, by obtaining the device information, the corresponding user information can be obtained, and the program directly determines the identity of the target user based on this user information. Or

[0093] Determine the target program currently in the running state, retrieve the user registration information of the target program, and determine the identity of the target user based on the registration information. For example, if the target program is a conference program, when the user wants to use the conference program, the user registers an identifier to distinguish from other users. When the user uses the conference program, the corresponding user identifier is displayed. When the user speaks, the voice of the user is collected, and the voice data is matched with the voiceprint data with the user identifier. The target program can be the program involved in the verification stage when the program determines the response to the voice data, or another program involved in determining the target program, which is not specified specifically. Or

[0094] Output a prompt message to the user, instructing the user to input identity information, and the program determines the identity of the target user based on the identity information input by the user. Or

[0095] Match the voice data with the voiceprint registration data stored during the local history period, and determine the identity of the target user based on the matching result.

[0096] The above various methods for determining the target user can be applied alternatively, or can be combined and applied in various ways. Moreover, the multiple determination methods in response to the voice data can also be applied alternatively, or can be combined and applied in various ways.

[0097] In one embodiment, as Figure 2 shown, the method further includes:

[0098] S6: In response to the acquisition of the voice data at the second moment, perform a first process or a second process on the voice data to obtain a first processing result or a second processing result corresponding to the second moment;

[0099] S7: Corresponding match the first processing result and the second processing result at the second moment with those at the first moment to obtain a first matching result;

[0100] S8: If the first matching result indicates that the similarity meets the requirements, perform voiceprint registration of the target user based on the first processing result or the second processing result at the second moment.

[0101] For example, after the voice data obtained at the first moment is processed to form a first processing result or a second processing result, the system can directly perform voiceprint registration based on this result, or can continue to receive the voice data at the next moment, and use the voice data at the next moment to verify the voice data at the previous moment, and finally use the verified voice data to extract voiceprint features to complete the subsequent voiceprint registration. Specifically, in response to the acquisition of the voice data meeting the conditions at the second moment, then perform the corresponding first process or second process on the voice data at the second moment to obtain a first processing result or a second processing result. Then compare the first processing result at the first moment with that at the second moment, or compare the second processing result at the first moment with that at the second moment to obtain a first matching result. If the first matching result indicates that the processing results corresponding to the two moments meet the similarity requirements, it means that the collected voice data is the voice data of the target user. At this time, the subsequent voiceprint registration of the target user can be performed based on the first processing result or the second processing result at the second moment. Or, the voiceprint registration of the target user can also be performed based on the voice data at the first moment and the second moment together. Or, the voiceprint registration of the target user can also be performed based on the voice data at the first moment.

[0102] The second moment of the above collection can be one moment or multiple moments, that is, the voice data of multiple moments can be collected, and then the voice data of the previous moment or any historical moment can be verified in turn by the processing and results of the voice data of the next moment. Finally, the voice data of the target user is determined, and the voiceprint registration of the target user is completed based on the processing results of the voice data. The first moment and the second moment can be different collection moments of the continuous voice data of the target user. For example, a speech of the target user in a meeting, or the collection moments of the voice data with time intervals of the target user, such as during multiple speeches of the target user.

[0103] Further, the method further includes:

[0104] S9: If the first matching result indicates that the similarity does not meet the requirement, obtain the voice data of the third moment, and perform the first processing or the second processing on the voice data of the third moment to obtain the first processing result or the second processing result corresponding to the third moment;

[0105] S10: Perform corresponding matching on the first processing results and the second processing results of the first moment, the second moment, and the third moment to obtain different second matching results;

[0106] S11: Perform voiceprint registration of the target user based on the first processing result or the second processing result of the two moments corresponding to the second matching result with the similarity meeting the requirement.

[0107] Continuing with the previous embodiment, if the first matching result indicates that the matching degree between the speech data at the first moment and the speech data at the second moment does not meet the requirements, it indicates that the speech data collected at the two moments may not all be input by the target user, and it is very likely to be input by other users. At this time, the program will continue to obtain the speech data at the third moment and perform the first processing or the second processing on the speech data at the third moment to obtain the corresponding first processing result or second processing result at the third moment. Then, the first processing results at the third moment, the second moment, and the first moment are compared, or the second processing results at the third moment, the second moment, and the first moment are compared. The comparison includes the comparison of the processing results of any two moments and the comparison of the processing results of the three moments, resulting in multiple different second matching results. After that, the program can determine the first processing result or the second processing result at two moments with a similarity that meets the requirements based on each matching relationship and the second matching result to perform voiceprint registration for the target user. That is, in this embodiment, when the matching result does not meet the requirements, indicating that the speech data collected at two moments does not correspond to the same user, the program continues to collect the speech data at the next moment and performs another comparison. This step can be repeated once or multiple times until the matching result meets the requirements and it is determined that the collected speech data is the speech data input by the target user. Then, based on the speech data of the corresponding target user, voiceprint registration for the target user is performed.

[0108] Suppose, in an online meeting scenario, the target user inputs the first speech data at the first moment, and a user in the environment inputs the second speech data at the second moment. At this time, the processing results of the speech data at the first moment and the second moment will not match because they correspond to the speech data of two different users. Therefore, the program will continue to collect the speech data at the third moment. Suppose the speech data at the third moment is the speech data input by the target user, then its processing result will match the processing result of the speech data at the first moment. At this time, the program will perform voiceprint registration for the target user based on the processing results of the speech data at the first moment and / or the third moment. On the contrary, if the speech data collected at the third moment is still the speech data input by the user in the environment, then continue to collect the speech data at the fourth moment, and so on, until the speech data of the target user is determined, and voiceprint registration for the target user is completed based on this.

[0109] In addition, for a target user who has already registered voiceprint data, the program can extract the historical voiceprint data and compare it with the processing result of the speech data at the first moment. If the matching degree does not meet the requirements, such as the matching degree does not exceed 80%, etc., it is determined that the voiceprint registration data needs to be updated or re-registered. At this time, the program will continue to collect the speech data at the next moment and compare it with the speech data at the previous moment until the new voiceprint data of the target user is finally determined. After obtaining the new voiceprint data, the program can update the historical voiceprint data or re-perform voiceprint registration for the target user.

[0110] Combined with Figure 3 As shown, in one embodiment, the first processing of the voice data includes:

[0111] S12: Determine the volume and duration of the voice data;

[0112] S13: When the volume and duration of the voice data meet the first requirement, determine the voiceprint data of the corresponding target user based on the voice data.

[0113] When only the voice characteristics of one user are included in the voice data, it indicates that the voice data is single-person voice data, that is, the voice data is the voice data of the target user. Therefore, the first processing needs to be performed on it to generate the first processing result. The first processing includes first determining whether the volume of the voice data meets the requirements, such as determining whether the volume meets the requirements through SPNR detection. If not, the voice data is discarded. If it meets the requirements, further determine whether the duration of the voice data meets the requirements. If it still meets the requirements, determine the voiceprint data of the corresponding target user based on the voice data.

[0114] Continue to combine Figure 3 As shown, the second processing of the voice data includes:

[0115] S14: Perform multi-person audio separation calculation on the voice data to extract the target voice data corresponding to the target user;

[0116] S15: Determine the volume and duration of the voice data;

[0117] S16: When the volume and duration of the voice data meet the first requirement, determine the voiceprint data of the corresponding target user based on the voice data.

[0118] In one embodiment, when the obtained voice data involves multiple users, that is, the collected voice data includes the voice data of multiple users at the same time, then multi-person audio separation calculation needs to be performed on the obtained voice data, such as dividing the voice data corresponding to each user, and then extracting and determining the voice data of the target user from the voice data of multiple users. Or directly extract the voice data of the target user from the voice data of multiple users. After determining the voice data of the target user, judge the volume and duration of the voice data. If the judgment results of the volume and duration meet the requirements, complete the voiceprint registration of the target user based on the voice data.

[0119] The multi-person audio separation calculation of the voice data to extract the target voice data corresponding to the target user includes:

[0120] S17: Determine the collection angles corresponding to the multiple pieces of collected voice data respectively;

[0121] S18: Extract the voice data corresponding to the target angle as the target voice data of the target user; or

[0122] S19: Determine the collection distances corresponding to the multiple pieces of collected voice data respectively;

[0123] S20: Extract the voice data corresponding to the target distance as the target voice data of the target user; or

[0124] S21: Determine the audio parameters corresponding to the multiple pieces of collected voice data respectively;

[0125] S22: Extract the voice data with the target audio parameters as the target voice data of the target user.

[0126] For example, in one embodiment, when collecting voice data, determine the collection angle of each piece of voice data. Different collection angles respectively correspond to a user. After obtaining the collection angles of each piece of voice data, divide the voice data with the same or similar collection angles into the voice data of a user. Accordingly, obtain the voice data of different users identified by the collection angle. Then, select from the voice data of multiple users the voice data of the user whose collection angle meets the requirements, that is, the collection angle is the target angle, and determine it as the target voice data of the target user, that is, the target voice data. The target angle can be but is not limited to an angle representing directly facing the device screen or an angle within a certain angle range.

[0127] Or, in another embodiment, when collecting voice data, determine the collection distance of each piece of voice data. Different collection distances represent different collection orientations of the voice data, and different collection orientations mean that different users in different orientations output voice data and are collected by the device. Therefore, in this way, it can be determined that the voice data with the collection distance meeting the similarity threshold belongs to the voice data output by the same user, and the collection distances corresponding to the voice data output by different users are different. Each piece of voice data can be identified by the user based on the collection distance. Then, select from the determined multiple pieces of voice data corresponding to different users the voice data corresponding to the target distance as the target voice data of the target user. The target distance can be the minimum distance, the vertical distance, etc., representing the user directly facing the device as the target user. Or, in yet another embodiment, the audio parameters of the multiple pieces of collected voice data can be determined, such as determining the volume, and the voice data with the target audio parameters can be used as the target voice data of the target user, such as using the voice data with the largest volume as the target voice data of the target user, etc.

[0128] If audio data of multiple people meeting the conditions are obtained, where the conditions are at least one of the acquisition angle, distance, and audio parameters. For example, if voice data of multiple people meeting the conditions are obtained at the first moment and voice data of multiple people meeting the conditions are obtained at the second moment, based on features of the voice such as timbre and pitch, the data of multiple people at the first moment and the data of multiple people at the second moment can be judged one by one to obtain the voice data at the first moment with a relatively high voice feature similarity and the corresponding voice data at the second moment for voiceprint registration.

[0129] Continuing to combine Figure 3 As shown, the voiceprint registration of the target user based on the first processing result or the second processing result includes:

[0130] S23: In response to an input instruction, perform voiceprint registration of the target user based on the first processing result or the second processing result; or

[0131] S24: In response to an input instruction, update the historical voiceprint registration data of the target user based on the first processing result or the second processing result.

[0132] After obtaining the voice data of the target user and processing it to obtain the first processing result or the second processing result, a prompt can be output to the user to determine whether to perform voiceprint registration. After receiving an input instruction indicating a definite registration, the program responds to this instruction and performs voiceprint registration of the target user based on the first processing result or the second processing result; or, the target user is a registered user. At this time, the program can output an indication to prompt to update the historical voiceprint data. When receiving an instruction for the target user to confirm the update, in response to this instruction, update the historical voiceprint registration data of the target user based on the first processing result or the second processing result.

[0133] This method can be executed on the terminal device side or in the cloud. If this method runs in the cloud, the cloud receives data sent by multiple terminal devices respectively. The data sent by each terminal device includes the identifier of the terminal device, the target user, and multiple pieces of collected voice data. For the data sent by each terminal device, the cloud processes them respectively, and can realize voiceprint registration of multiple target users when multiple terminal devices open the same application program or the same function.

[0134] As Figure 4 shown, another embodiment of the present application simultaneously provides a voice data processing device 100, including:

[0135] A first response module, configured to determine a target user in response to the acquisition of voice data at the first moment, and identify the matching relationship between the voice data and the target user;

[0136] A second response module, configured to determine whether the voice data is single-person voice data or multi-person voice data in response to the recognition result being the first recognition result;

[0137] A first processing module, configured to perform a first processing on the voice data when it is single-person voice data;

[0138] A second processing module, configured to perform a second processing on the voice data when it is multi-person voice data, where the second processing is different from the first processing, and both the first processing and the second processing are used to obtain the voiceprint data corresponding to the target user;

[0139] A registration module, configured to perform voiceprint registration of the target user based on the first processing result or the second processing result.

[0140] In one embodiment, the apparatus further includes:

[0141] A third response module, configured to perform a first processing or a second processing on the voice data in response to obtaining the voice data at a second moment, to obtain a first processing result or a second processing result corresponding to the second moment;

[0142] A first matching module, configured to perform corresponding matching on the first processing results and the second processing results at the second moment and the first moment to obtain a first matching result;

[0143] The registration module is further configured to perform voiceprint registration of the target user based on the first processing result or the second processing result at the second moment when the first matching result indicates that the similarity meets the requirements.

[0144] In one embodiment, the apparatus further includes:

[0145] An acquisition module, configured to acquire the voice data at a third moment and perform a first processing or a second processing on the voice data at the third moment to obtain a first processing result or a second processing result corresponding to the third moment when the first matching result indicates that the similarity does not meet the requirements;

[0146] A second matching module, configured to perform corresponding matching on the first processing results and the second processing results at the first moment, the second moment, and the third moment to obtain different second matching results;

[0147] The registration module is further configured to perform voiceprint registration of the target user based on the first processing result or the second processing result at two moments corresponding to the second matching result whose similarity meets the requirements.

[0148] In one embodiment, determining the target user includes:

[0149] Obtaining device information and determining the target user based on the device information; or

[0150] Obtain the registration information of the target program, and determine the target user based on the registration information; or

[0151] Determine the target user based on the input identity information; or

[0152] Determine the target user based on the stored historical voiceprint registration data.

[0153] In one embodiment, the obtaining of the voice data in response to the first moment includes:

[0154] Obtain the voice data within the specified sound frequency range at the first moment; or

[0155] Obtain the voice data whose volume meets the requirements at the first moment; or

[0156] When it is determined that there is audio input and output, obtain the voice data at the first moment; or

[0157] When it is determined that the target function is started, obtain the voice data at the first moment; or

[0158] When it is determined that the target program is started, obtain the voice data at the first moment.

[0159] In one embodiment, the first processing of the voice data includes:

[0160] Determine the volume and duration of the voice data;

[0161] When the volume and duration of the voice data meet the first requirements, determine the voiceprint data corresponding to the target user based on the voice data.

[0162] In one embodiment, the second processing of the voice data includes:

[0163] Perform multi-person audio separation calculation on the voice data, and extract the target voice data corresponding to the target user;

[0164] Determine the volume and duration of the voice data;

[0165] When the volume and duration of the voice data meet the first requirements, determine the voiceprint data corresponding to the target user based on the voice data.

[0166] In one embodiment, the performing multi-person audio separation calculation on the voice data and extracting the target voice data corresponding to the target user includes:

[0167] Determine the collection angles corresponding to the collected multiple voice data respectively;

[0168] Extract the voice data corresponding to the target angle as the target voice data of the target user; or

[0169] Determine the collection distances corresponding to the multiple collected voice data respectively;

[0170] Extract the voice data corresponding to the target distance as the target voice data of the target user

[0171] Determine the audio parameters corresponding to the multiple collected voice data respectively;

[0172] Extract the voice data with the target audio parameters as the target voice data of the target user.

[0173] In one embodiment, the voiceprint registration of the target user based on the first processing result or the second processing result includes:

[0174] In response to an input instruction, perform voiceprint registration of the target user based on the first processing result or the second processing result; or

[0175] In response to an input instruction, update the historical voiceprint registration data of the target user based on the first processing result or the second processing result.

[0176] Another embodiment of the present application further provides an electronic device, including;

[0177] One or more processors;

[0178] A memory configured to store one or more programs;

[0179] When the one or more programs are executed by the one or more processors, the one or more processors implement the voice data processing method described in any one of the above.

[0180] Furthermore, an embodiment of the present application further provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the voice data processing method described above is implemented. It should be understood that each solution in this embodiment has the corresponding technical effects in the above method embodiment, and will not be elaborated here.

[0181] Furthermore, an embodiment of the present application further provides a computer program product, the computer program product is tangibly stored on a computer-readable medium and includes computer-readable instructions, and when the computer-executable instructions are executed, at least one processor is caused to execute the voice data processing method in the above-described embodiment.

[0182] It should be noted that the computer storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable medium can, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access storage medium (RAM), a read-only storage medium (ROM), an erasable programmable read-only storage medium (EPROM or flash memory), an optical fiber, a portable compact disk read-only storage medium (CD-ROM), an optical storage medium, a magnetic storage medium, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program configured to be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, antenna, optical cable, RF, etc., or any suitable combination of the above.

[0183] In addition, those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0184] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing the process Figure 1one or more processes and / or blocks Figure 1 a system for the functions specified in one or more blocks

[0185] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction system that implements the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 in one or more blocks

[0186] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is exemplary only and is not intended to imply that the scope of the present application is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of brevity.

Claims

1. A method for processing speech data, comprising: In response to obtaining the voice data at the first moment, determining a target user, and identifying a matching relationship between the voice data and the target user; In response to the recognition result being the first recognition result, determining that the voice data is single-person voice data or multi-person voice data; When the voice data is from a single person, performing a first processing on the voice data; When the voice data is multi-person voice data, performing a second processing on the voice data, the second processing is different from the first processing, and both the first processing and the second processing are used to obtain the voiceprint data corresponding to the target user; The voiceprint registration of the target user is performed based on the first processing result or the second processing result.

2. The method for processing speech data according to claim 1, wherein: The method further comprises: In response to obtaining the voice data at the second moment, performing a first process or a second process on the voice data to obtain a first process result or a second process result corresponding to the second moment; Matching the first processing result and the second processing result at the second moment with the first moment to obtain a first matching result; If the first matching result represents that the similarity meets the requirement, the voiceprint registration of the target user is performed based on the first processing result or the second processing result at the second moment.

3. The method for processing speech data according to claim 2, wherein: The method further comprises: If the first matching result indicates that the similarity does not meet the requirement, acquiring the voice data at the third moment, and performing the first processing or the second processing on the voice data at the third moment to obtain the first processing result or the second processing result corresponding to the third moment; Matching the first processing result and the second processing result at the first moment, the second moment, and the third moment accordingly to obtain different second matching results; The voiceprint registration of the target user is performed based on the first processing results or the second processing results at two moments corresponding to the second matching results that meet the similarity requirements.

4. The method for processing speech data according to claim 1, wherein: The determining of the target user comprises: obtaining device information, and determining the target user based on the device information; or Obtaining registration information of a target program, and determining the target user based on the registration information; or Determine the target user based on the identity information entered; or Determine the target user based on the stored historical voiceprint registration data.

5. The method for processing speech data according to claim 1, wherein: The acquisition of the voice data in response to the first moment includes: In response to obtaining speech data within a specified sound frequency range at a first moment; or In response to obtaining voice data whose volume satisfies the requirement at the first moment; or When determining that there is audio input and output, responding to obtaining voice data at a first moment; or When determining that the target function is activated, in response to obtaining the voice data at the first moment; or When it is determined that the target program is started, the acquisition of voice data at a first moment is responded to.

6. The method for processing speech data according to claim 1, wherein: The first processing of the voice data includes: Determining the volume and duration of the voice data; When the volume and duration of the voice data meet the first requirement, the voiceprint data of the corresponding target user is determined based on the voice data.

7. The method for processing speech data according to claim 1, wherein: Performing a second process on the voice data includes: Performing multi-person audio separation calculation on the voice data to extract target voice data corresponding to the target user; Determining the volume and duration of the voice data; When the volume and duration of the voice data meet the first requirement, the voiceprint data of the corresponding target user is determined based on the voice data.

8. The method for processing speech data according to claim 9, wherein: The performing multi-person audio separation calculation on the voice data to extract target voice data corresponding to the target user includes: Determine the collection angles corresponding to the collected multiple voice data respectively; Extracting the speech data corresponding to the target angle as the target speech data of the target user; or Determine the collection distances corresponding to the collected multiple voice data respectively; Extracting the voice data corresponding to the target distance as the target voice data of the target user; or Determine the audio parameters corresponding to the collected multiple voice data respectively; The extracted voice data having the target audio parameters is the target voice data of the target user.

9. The method for processing speech data according to claim 1, wherein: The registering the voiceprint of the target user based on the first processing result or the second processing result includes: In response to an input instruction, registering the target user's voiceprint based on the first processing result or the second processing result; or In response to the input instruction, the historical voiceprint registration data of the target user is updated based on the first processing result or the second processing result.

10. A voice data processing device, comprising: A first response module, configured to determine a target user in response to obtaining the voice data at a first moment, and identify a matching relationship between the voice data and the target user; A second response module, configured to determine, in response to the recognition result being the first recognition result, that the voice data is single-person voice data or multi-person voice data; A first processing module, used for performing a first processing on the voice data when the voice data is single-person voice data; A second processing module, for performing a second processing on the voice data when the voice data is multi-person voice data, wherein the second processing is different from the first processing, and both the first processing and the second processing are used to obtain the voiceprint data corresponding to the target user; A registration module is used to register the voiceprint of the target user based on the first processing result or the second processing result.