Voiceprint noise reduction method and device, call device, and storage medium
By copying voiceprint features to ADSP and initializing the voice algorithm before the calling device starts up, the problem of noise interference in VoIP calls is solved, and the clarity and quality of voice calls are improved efficiently.
Patent Information
- Application Number
- CN202110898331.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-05
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-08-05
AI Technical Summary
In VoIP calls, significant noise interference results in low voice clarity, and current technologies cannot effectively reduce the impact of noise interference, leading to a decline in voice call quality.
Before the call device is started, the first voiceprint feature on the application side is copied to the audio digital signal processor module (ADSP) via shared memory. During the call, the voice algorithm is initialized, and the voiceprint feature is extracted and matched by the ADSP to generate voice call data.
It reduces the impact of voiceprint noise reduction on call latency, improves the clarity and quality of voice calls, and provides a better user experience.
Smart Images

Figure CN115706746B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of voiceprint noise reduction, and particularly relates to a voiceprint noise reduction method and device, a call device, and a storage medium. BACKGROUND
[0002] In a VoIP (Voice over Internet Protocol) handheld mode, if there is a large noise interference in the current call environment, for example, other people talking, large wind noise, car horn sound, etc., the audio component will collect the interference voiceprint during the call process, forming a voiceprint noise, resulting in low voice call clarity, and thus reducing the voice call quality. SUMMARY
[0003] To overcome the problems in the related art, the present disclosure provides a voiceprint noise reduction method and device, a call device, and a storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, a voiceprint noise reduction method is provided, applied to a call device, and including:
[0005] Before the call device starts a call, copying a first voiceprint feature pre-stored by an application side of the call device to an audio digital signal processor module (ADSP) of the call device through a shared memory;
[0006] During the call process of the call device, in response to a noise reduction enabling operation on a call interface, initializing a voice algorithm in the ADSP, wherein the initialized voice algorithm can establish communication between the ADSP and the application side; and
[0007] Transmitting sound data in the call process from an audio hardware layer of the call device to the ADSP based on the communication, so that the ADSP extracts a second voiceprint feature of the sound data based on the initialized voice algorithm;
[0008] Determining a target voiceprint feature matching the first voiceprint feature from the second voiceprint feature based on the initialized voice algorithm in the ADSP;
[0009] Generating voice call data according to the target voiceprint feature.
[0010] According to a second aspect of an embodiment of the present disclosure, a voiceprint noise reduction device is provided, and the device includes:
[0011] a copying module configured to copy, by means of shared memory, first voiceprint features pre-stored on an application side of the call device to an audio digital signal processor module ADSP of the call device before the call device starts a call;
[0012] an initializing module configured to initialize a voice algorithm in the ADSP in response to a noise reduction enabling operation on a call interface during a call process of the call device, wherein the initialized voice algorithm can establish communication between the ADSP and the application side; and
[0013] a transmitting module configured to transmit, from an audio hardware layer of the call device, sound data in the call process to the ADSP based on the communication, so that the ADSP extracts second voiceprint features of the sound data based on the initialized voice algorithm;
[0014] a determining module configured to determine, based on the initialized voice algorithm in the ADSP, target voiceprint features matching the first voiceprint features from the second voiceprint features;
[0015] a generating module configured to generate voice call data according to the target voiceprint features.
[0016] According to a third aspect of the embodiments of the present disclosure, a call device is provided, comprising:
[0017] a processor;
[0018] a memory for storing processor-executable instructions;
[0019] wherein the processor is configured to:
[0020] copy, by means of shared memory, first voiceprint features pre-stored on an application side of the call device to an audio digital signal processor module ADSP of the call device before the call device starts a call;
[0021] initialize a voice algorithm in the ADSP in response to a noise reduction enabling operation on a call interface during a call process of the call device, wherein the initialized voice algorithm can establish communication between the ADSP and the application side; and
[0022] transmit, from an audio hardware layer of the call device, sound data in the call process to the ADSP based on the communication, so that the ADSP extracts second voiceprint features of the sound data based on the initialized voice algorithm;
[0023] determine, based on the initialized voice algorithm in the ADSP, a target voiceprint feature matching the first voiceprint feature from the second voiceprint feature;
[0024] generate voice call data according to the target voiceprint feature.
[0025] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer program instructions. When the computer program instructions are executed by a processor, the steps of the method in any one of the first aspect are implemented.
[0026] The technical solutions provided by the embodiments of the present disclosure can have the following beneficial effects: by copying the first voiceprint feature pre-stored by the application side of the call device to the audio digital signal processor module ADSP of the call device through shared memory before the call device starts a call; in the call process of the call device, in response to a noise reduction enabling operation on the call interface, a voice algorithm in the ADSP is initialized, the initialized voice algorithm can establish communication between the ADSP and the application side; sound data in the call process is transmitted from the audio hardware layer of the call device to the ADSP based on the communication, so that the ADSP extracts a second voiceprint feature of the sound data based on the initialized voice algorithm; based on the initialized voice algorithm in the ADSP, a target voiceprint feature matching the first voiceprint feature is determined from the second voiceprint feature; and voice call data is generated according to the target voiceprint feature. By copying the first voiceprint feature from the application side to the ADSP through shared memory before the call is started, the voice algorithm of the ADSP can directly use the first voiceprint feature after the voice algorithm of the ADSP is initialized, without waiting for the first voiceprint feature to be obtained from the application side during the call process, thereby reducing the impact of voiceprint noise reduction on call delay.
[0027] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0028] The accompanying drawings, which are incorporated into and form part of the specification, illustrate one embodiment consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0029] Figure 1 is a flowchart of a voiceprint noise reduction method according to an exemplary embodiment.
[0030] Figure 2 is a flowchart of a method for implementing Figure 1 step S11 in the method.
[0031] Figure 3 is a schematic diagram of a path for copying a first voiceprint feature according to an exemplary embodiment.
[0032] Figure 4 This is a schematic diagram of a control panel for a web calling assistant according to an exemplary embodiment.
[0033] Figure 5 This is a schematic diagram of a control panel for recording voiceprints according to an exemplary embodiment.
[0034] Figure 6 This is a schematic diagram illustrating a control panel for recording text of various types according to an exemplary embodiment.
[0035] Figure 7 This is a schematic diagram of a control panel for recording voiceprints according to an exemplary embodiment.
[0036] Figure 8 This is a schematic diagram of a control panel for a voiceprint noise reduction function according to an exemplary embodiment.
[0037] Figure 9 This is a schematic diagram illustrating the setting path of a voiceprint noise reduction function according to an exemplary embodiment.
[0038] Figure 10 This is a schematic diagram illustrating code for storing a first voiceprint feature in the kernel according to an exemplary embodiment.
[0039] Figure 11 This is a schematic diagram illustrating the retrieval of voiceprint data during a call, according to an exemplary embodiment.
[0040] Figure 12 This is an implementation illustrated according to an exemplary embodiment. Figure 1 The flowchart for step S13.
[0041] Figure 13 This is a flowchart illustrating an embodiment of obtaining a first voiceprint feature.
[0042] Figure 14 This is a schematic diagram illustrating a method of using voiceprint data for model training according to an exemplary embodiment.
[0043] Figure 15 This is a block diagram illustrating a voiceprint noise reduction device 100 according to an exemplary embodiment.
[0044] Figure 16 This is a block diagram illustrating an apparatus 800 for voiceprint noise reduction according to an exemplary embodiment. Detailed Implementation
[0045] The exemplary embodiments will be described in detail herein below with reference to the drawings. In the following description, unless otherwise indicated, like numbers in the different drawings represent similar or identical elements. The following exemplary embodiments described are not meant to represent all implementations consistent with the present disclosure. Rather, they are simply examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0046] Figure 1 is a flow chart of a voiceprint noise reduction method according to an exemplary embodiment, as Figure 1 shown, the method is used in a call device, comprising the following steps.
[0047] In step S11, before the call device starts a call, the first voiceprint feature pre-stored by the application side of the call device is copied to the audio digital signal processor module ADSP of the call device through shared memory.
[0048] Wherein, before the call device starts a call, it can be after the system of the call device is powered on, or after the first voiceprint feature is input.
[0049] On the basis of the above-mentioned embodiments, the first voiceprint feature is pre-stored in the kernel of the application side, Figure 2 is a flow chart of a method for implementing Figure 1 step S11 according to an exemplary embodiment, as Figure 2 shown, step S11 comprises the following steps.
[0050] In step S111, a shared memory area between the kernel and the audio digital signal processor module ADSP is created.
[0051] Wherein, communication can be carried out between the application side and the driver side through a communication control, and the application side sends a configuration command to the driver side, and the driver side creates a shared memory area after receiving the configuration command. The configuration command is:
[0052] ctl = mixer_get_ctl_by_name(adev->mixer, "Qdsp voice model Data")
[0053] mixer_ctl_set_array(ctl, (void*)&modelinfo, sizeof(struct model_info))
[0054] As Figure 3As shown, after receiving the configuration command, the drive side applies to create a shared memory region (Share Buffer1) between the application side AP and the ADSP, and after creating the shared memory region, enables the application side to copy the first voiceprint feature pre-stored in the kernel to the shared memory region.
[0055] In step S112, the first voiceprint feature pre-stored in the kernel is copied to the shared memory region.
[0056] Wherein, the first voiceprint feature is copied to the shared memory region through the following message ID: #define AFE_PORT_SEND_DATA_CMD 0x00011111.
[0057] In step S113, after the first voiceprint feature is copied to the shared memory region, the first voiceprint feature in the shared memory region is copied to the static service buffer of the ADSP.
[0058] Wherein, the static service of the ADSP receives the message command from the front-end interface indicating that the copying of the first voiceprint feature to the shared memory region is completed through the message processing function afe_svc_apr_msg_handler(), and then Figure 3 As shown, based on the static service (AFE staticservice) of the ADSP, the first voiceprint feature in the shared memory region is copied to the static service buffer (Share Buffer2) of the ADSP for use by the voice algorithm in the noise cancellation module (ECNS dynamicmodule) after initialization is completed.
[0059] On the basis of the above embodiment, in step S113, after the first voiceprint feature is copied to the shared memory region, the first voiceprint feature in the shared memory region is copied to the static service buffer of the ADSP based on the out-of-band message mechanism OOB (out-of-band).
[0060] Specifically, a two-end connection of out-of-band data between the shared memory region and the static service of the ADSP is established, and if the copying of the first voiceprint feature in the kernel is completed, the static service of the ADSP can be immediately notified, instead of being notified through the message queuing mechanism, and in the case of a large amount of data of the first voiceprint feature, the first voiceprint feature in the shared memory region can also be quickly copied to the static service buffer of the ADSP.
[0061] The first voiceprint feature pre-stored can be copied into the static service buffer of the ADSP through the shared memory and the out-of-band message mechanism, the speed of copying the first voiceprint feature into the static service buffer of the ADSP is improved, and the voiceprint noise reduction is avoided from affecting the call delay.
[0062] In step S12, in the call process of the call device, the voice algorithm in the ADSP is initialized in response to the noise reduction enabling operation on the call interface, and the initialized voice algorithm can establish communication between the ADSP and the application side.
[0063] The application side communicates with the ADSP through audio management parameters (AudioManager setParameters) and audio system parameters (AudioRecord setParameters).
[0064] In an embodiment, in response to the noise reduction enabling operation on the call interface, if the ADSP fails to obtain the first voiceprint feature, the first voiceprint feature is prompted to be input. As shown in Figure 4 and 5 As shown, the application program that can start the voiceprint noise reduction function and the corresponding switch button are displayed, and in the case that the voiceprint noise reduction function registers the voiceprint start, a prompt window of "your voiceprint is needed to start this function" is popped up at the bottom of the call interface, and the selection buttons of "cancel" and "record immediately" are displayed.
[0065] In an embodiment, as shown in Figure 4 In the popped-up network phone assistant settings, the application program that needs to start the voiceprint recording and the voiceprint noise reduction can be displayed, and the voiceprint recording of the corresponding application program is started by the user manually. After the voiceprint data is recorded successfully for the first time, the voiceprint recording interface can be displayed as shown in Figure 4 As shown, the option of re-recording the voiceprint under the voiceprint noise reduction function switch is displayed.
[0066] In an embodiment, as shown in Figure 6 When the voiceprint is recorded, the recording text of different voice scene types can be selected, for example, one or more types of recording text in the story text, the earthy love, the movie lines, and the simple text. After the recording text of the corresponding type is selected, as shown in Figure 7 The corresponding voiceprint recording user interface is popped up in response to the user selecting the recording text of the corresponding type and touching "start recording", and the voiceprint recording interface corresponding to the recording text is popped up.
[0067] In an embodiment, the voiceprint noise reduction function is in a closed state when the call device is in a call, i.e., in this embodiment, the voiceprint noise reduction function is by default in a closed state, so as to reduce the resource occupation of the voiceprint noise reduction on the call device, and to avoid sending various key-value pairs and control instructions when the call is opened, thereby reducing the signaling overhead.
[0068] Specifically, the ADSP records the switch state of the voiceprint noise reduction switch through "data[0]:indicate spkrid recognition enable / disable". The application side closes or opens the voiceprint noise reduction function by sending the opening key-value pair "Audio Manager.setParameters("spkrid_recognition=on / off")" to the audio hardware communication layer (Audio HAL).
[0069] As shown in Figure 8 , in the case where the mode of the call is a network call mode, the network call assistant user interface as shown in Figure 8 may be popped up in response to a touch on the network call assistant pop-up window in the upper left corner of the call interface, and a control panel for opening and closing the voiceprint noise reduction function is displayed, and then in response to a selection of the voiceprint noise reduction enable operation, the opening key-value pair is sent to indicate that the voiceprint noise reduction function is opened, and the voice algorithm in the ADSP is initialized.
[0070] In an embodiment, in the case where the call is ended, the voiceprint noise reduction function of the application side is by default switched to a closed state. Here, the application side can refer to an application program, i.e., in the case where the call is ended, the voiceprint noise reduction function of the application program is by default switched to a closed state.
[0071] In a possible scenario, if the application side has no permission to enable the voiceprint noise reduction, a prompt window is used to show a setting path for enabling the voiceprint noise reduction permission of the corresponding application side, so as to enable the voiceprint noise reduction function according to the setting path. As shown in Figure 9 , the user is guided to open the voiceprint noise reduction function through the following setting path: Settings->Features->Network Call Assistant->Voiceprint Noise Reduction. In the case where the network call assistant is not selected, the voiceprint noise reduction selection switch is in a hidden state, and in the case where the network call assistant is selected, the voiceprint noise reduction selection switch is switched to a display state through a drop-down menu.
[0072] As can be understood, as shown in Figure 9 , each application program has a switch for opening or closing the voiceprint noise reduction function, and the voiceprint noise reduction function of different application programs can be opened or closed respectively.
[0073] After the language algorithm is initialized, the ADSP can establish communication with the application side based on the audio calibration (Runtime Audio Calibration) mode.
[0074] In step S13, the sound data in the call process is transmitted from the audio hardware layer of the call device to the ADSP based on the communication, so that the ADSP extracts the second voiceprint feature of the sound data based on the initialized voice algorithm.
[0075] In the call process, if the communication between the audio calibration mode and the application side is successful, the sound data in the call process is transmitted from the audio hardware layer of the call device to the ADSP based on the communication, and the initialized voice algorithm in the ADSP can extract the voiceprint feature in the sound data, for example, the timbre, voiceprint, etc. in the voiceprint data.
[0076] Further, the extracted voiceprint feature input vector extractor obtains the sound data corresponding voiceprint vector, and the voiceprint vector is taken as the second voiceprint feature.
[0077] Among them, the length of the voiceprint data is recorded by "data[1,...]:store voice model data len". The length of the voiceprint data can be used to represent the comparison length of the first voiceprint feature and the second voiceprint feature in the call process. For example, the second voiceprint feature of 1024 bytes is compared with the first voiceprint feature.
[0078] In step S14, the target voiceprint feature matched with the first voiceprint feature is determined from the second voiceprint feature based on the initialized voice algorithm in the ADSP.
[0079] Among them, the ADSP compares the first voiceprint feature obtained from the static service buffer with the second voiceprint feature based on the initialized voice algorithm, and determines the target voiceprint feature matched with the first voiceprint feature from the second voiceprint feature.
[0080] In step S15, the voice call data is generated according to the target voiceprint feature.
[0081] It can be explained that the method provided by the present disclosure can be applied to the fields of voiceprint unlocking, voiceprint recognition, voice call and game voice, etc. Therefore, the voice call data can be used for voiceprint unlocking, voiceprint recognition, sending to the opposite end to meet the voice call and game voice, which can reduce the influence of noise such as environmental noise, human speaking sound, etc. on voice data collection, and improve the accuracy of voice data.
[0082] During the call, non-target voiceprint features such as environmental noise and speech of non-target personnel are filtered out, and the generated voice call data only contains target voiceprint features, so that the non-target voiceprint features cannot be transmitted to the call opposite end, and the influence of noise on voice can be reduced.
[0083] By using the technical solution, the first voiceprint feature is copied from the application side to the ADSP through the shared memory before the call is started. After the voice algorithm of the ADSP is initialized, the first voiceprint feature can be directly used, and there is no need to wait for acquisition from the application side during the call, so that the influence of voiceprint noise reduction on call delay is reduced. The user can hardly perceive the call delay, and the use experience of voiceprint noise reduction can be improved.
[0084] On the basis of the above-mentioned embodiment, the first voiceprint feature is stored in the static buffer in a pre-defined storage manner, wherein the pre-defined storage manner comprises storing the first voiceprint feature in a 32-bit compilation manner and a 64-bit compilation manner respectively.
[0085] As shown in Figure 10 , by defining the byte, the compatibility of the voiceprint data between the kernel and the ADSP is realized, so that the copying between the kernel and the ADSP can be quickly completed.
[0086] On the basis of the above-mentioned embodiment, in step S13, the sound data during the call is transmitted from the audio hardware layer of the call device to the ADSP based on the communication, comprising:
[0087] In a case where it is determined that the mode of the call is a network call mode, the sound data is acquired from the audio path of the audio hardware communication layer of the call device based on the application side, and the sound data is transmitted to the ADSP through the audio calibration communication between the application side and the ADSP.
[0088] As shown in Figure 11 , in a network call mode (Voice over Internet Protocol, VoIP), an audio path is used, and the sound data can be transmitted to the ADSP by directly calling an audio calibration communication “ / acdb_set_audio_cal()” interface in a real-time audio calibration (Runtime Audio Calibration) manner. The ADSP performs feature extraction and other preprocessing on the sound data, and based on a voice algorithm module, the first voiceprint feature is called from the static service buffer of the ADSP, and the target voiceprint feature matched with the first voiceprint feature is determined from the second voiceprint feature.
[0089] The technical solution can directly call the audio calibration communication interface in the VoIP scenario, and then quickly call the voiceprint data from the audio hardware communication layer to the ADSP, so that the user can hardly feel the delay of the voice.
[0090] On the basis of the above-mentioned embodiment, Figure 12 is a flowchart for obtaining a first voiceprint feature according to an example embodiment, as shown in Figure 1 The flowchart of step S13 is shown in Figure 12 In step S13, the audio hardware layer of the call device transmits the sound data during the call to the ADSP based on the communication, including:
[0091] In step S131, in the case of determining that the mode of the call is a direct outbound dialing mode, the call log of the debug interface based on the preset audio calibration tool is called, and the target call interface of the audio hardware communication layer of the call device is determined from the debug interface through the application side.
[0092] It can be explained that in the prior art, in the direct outbound dialing mode, a voice path is used, and because the audio calibration communication interface only supports ADSP_ASM_SERVICE and ADSP_ADM_SERVICE services, it cannot support dynamic configuration of the voice path.
[0093] Based on the above reasons, the call log of the debug interface based on the preset audio calibration tool QACT can be called to create a voice calibration interface "acdb_set_voice_cal()" to support the processing flow management and voice flow management in the direct outbound dialing mode.
[0094] In step S132, the sound data is obtained from the audio hardware communication layer according to the target call interface.
[0095] In step S133, the sound data is transmitted to the ADSP through the audio calibration communication between the application side and the ADSP.
[0096] With the above technical solution, the voice data can be successfully transmitted to the ADSP in the direct outbound dialing mode, and voiceprint noise reduction in the direct outbound dialing mode can be realized.
[0097] On the basis of the above-mentioned embodiment, Figure 13 is a flowchart for obtaining a first voiceprint feature according to an example embodiment, as shown in Figure 13 The flowchart includes the following steps.
[0098] In step S21, in response to the voiceprint recording operation on the settings interface of the call device, the audio hardware is invoked through a preset call key value, which is used to instruct the audio hardware to create a recording path using a specified recording device.
[0099] In response to the user's voiceprint recording operation for the application in the settings interface, the application sends a preset key-value pair "Audio Record.setParameters("mic_mode=spkrid_enroll")" to the audio hardware to ensure that a specific recording algorithm ID (acdbID=10029) is called to avoid calling audio hardware that is not suitable for voiceprint recording.
[0100] Upon receiving the call key-value pair, the audio hardware creates a recording path using the recording device specified in the call key-value pair, and then collects sound data through the specified recording device based on the recording path.
[0101] In step S22, sound data is collected using the recording device.
[0102] In step S23, a preset loading key-value pair is used to instruct the audio hardware communication layer of the calling device to load the sound data collected by the recording device into the ADSP in the calling device, so that the ADSP can extract the first voiceprint feature of the sound data collected by the recording device.
[0103] like Figure 14 As shown, after the voiceprint data acquisition is completed, the ADSP sends a preset loading key-value pair "AudioManager.setParameters("spkrid_model_update=true")" to the audio hardware communication layer. The audio hardware communication layer instructs the audio front-end (Audio Record) and trains the general background model (GenSpkID) in the application (Process) on the application side, and loads the sound data acquired by the recording device into the ADSP in the call device.
[0104] In one implementation, the validity of the recorded voiceprint data is determined based on the length of the voiceprint data or the duration of the voiceprint data recording. For example, if the length of the voiceprint data does not reach the preset data length, or the duration of the voiceprint data recording does not reach the preset time length, the voiceprint data recording is deemed to have failed.
[0105] Furthermore, the voiceprint data is verified based on the voice samples, and the voiceprint data is characterized based on the rule model and knowledge model to determine the authenticity of the collected voiceprint data. Then, the first voiceprint feature is determined based on the voice samples of closed-level particles.
[0106] Optionally, in step S23, the ADSP extracts the first voiceprint features of the sound data collected by the recording device, including:
[0107] training a general background model according to the sound data collected by the recording device to obtain a trained voiceprint recognition model, and the first voiceprint features of the sound data collected by the recording device.
[0108] Wherein, by mapping the voiceprint data to a low-dimensional space to maximize the data distance between classes and minimize the data distance within classes, a voiceprint recognition model PLDA and a voiceprint extractor are trained based on a general background model (GenSpkID) and a general extractor. The general background model is obtained by preprocessing a large-scale speaker voice and using the voiceprint features after extracting the voiceprint features as training samples.
[0109] and sequentially extracting features from the input voiceprint data, and extracting feature vectors from the voiceprint data after feature extraction by the voiceprint extractor to obtain first voiceprint features, and saving the voiceprint recognition model PLDA, the voiceprint extractor and the first voiceprint features according to the saving path " / data / vendor / model" as shown in Figure 14
[0110] In step S14, the voice algorithm in the ADSP based on the initialization is completed to determine the target voiceprint features matching the first voiceprint features from the second voiceprint features, including:
[0111] inputting the first voiceprint features and the second voiceprint features into the trained voiceprint recognition model to obtain the target voiceprint features in the second voiceprint features matching the first voiceprint features.
[0112] Wherein, the first voiceprint features and the second voiceprint features are input into the pre-trained voiceprint recognition model to obtain the matching score result of the second voiceprint features output by the voiceprint recognition model and the first voiceprint features, and the second voiceprint features meeting the preset threshold are taken as the target voiceprint features matching the first voiceprint features according to the matching score result.
[0113] Based on the same inventive concept, the present disclosure also provides a voiceprint noise reduction device, which can realize all or part of the steps of the voiceprint noise reduction method in the form of software, hardware or a combination of both. Figure 15 is a block diagram of a voiceprint noise reduction device 100 according to an exemplary embodiment, as shown in Figure 15 As shown, the device 100 comprises a copying module 110, an initializing module 120, a transmission module 130, a determining module 140 and a generating module 150.
[0114] The copying module 110 is configured to copy, by means of shared memory, a first voiceprint feature pre-stored on an application side of the call device to an audio digital signal processor module ADSP of the call device before the call device starts a call.
[0115] The initializing module 120 is configured to initialize a voice algorithm in the ADSP in response to a noise reduction enabling operation on a call interface during a call process of the call device, wherein the initialized voice algorithm can establish communication between the ADSP and the application side; and
[0116] The transmission module 130 is configured to transmit, from an audio hardware layer of the call device, sound data in the call process to the ADSP based on the communication, so that the ADSP extracts a second voiceprint feature of the sound data based on the initialized voice algorithm.
[0117] The determining module 140 is configured to determine, based on the initialized voice algorithm in the ADSP, a target voiceprint feature matching the first voiceprint feature from the second voiceprint feature.
[0118] The generating module 150 is configured to generate voice call data according to the target voiceprint feature.
[0119] The device described above can directly use the first voiceprint feature after the voice algorithm of the ADSP is initialized, without waiting for acquisition from the application side, thereby reducing the impact of voiceprint noise reduction on call delay.
[0120] Optionally, the first voiceprint feature is pre-stored in a kernel of the application side, and the copying module 110 is configured to create a shared memory area between the kernel and an audio digital signal processor module ADSP.
[0121] The first voiceprint feature pre-stored in the kernel is copied into the shared memory area.
[0122] After the first voiceprint feature is copied into the shared memory area, the first voiceprint feature in the shared memory area is copied into a static service buffer of the ADSP.
[0123] Optionally, the copying module 110 is configured to copy the first voiceprint feature in the shared memory region into a static service buffer of the ADSP based on an out-of-band message mechanism after the copying of the first voiceprint feature to the shared memory region is completed.
[0124] Optionally, the first voiceprint feature is stored in the static buffer in a predefined storage manner, wherein the predefined storage manner comprises storing the first voiceprint feature in a 32-bit compilation manner and a 64-bit compilation manner respectively.
[0125] Optionally, the transmission module 130 is configured to, in a case where it is determined that the mode of the call is a network call mode, acquire the sound data based on the application side from an audio path of an audio hardware communication layer of the call device, and transmit the sound data to the ADSP through audio calibration communication between the application side and the ADSP.
[0126] Optionally, the transmission module 130 is configured to, in a case where it is determined that the mode of the call is a direct external dialing mode, determine a target call interface of the audio hardware communication layer of the call device from a debugging interface based on a call log of a preset audio calibration tool call debugging interface;
[0127] acquire the sound data from the audio hardware communication layer according to the target call interface;
[0128] transmit the sound data to the ADSP through audio calibration communication between the application side and the ADSP.
[0129] Optionally, the copying module 110 is configured to:
[0130] in response to a voiceprint recording operation on a setting interface of the call device, call an audio hardware through a preset call key value pair, the call key value pair being used to instruct the audio hardware to create a recording path using a specified recording device;
[0131] acquire sound data according to the recording device;
[0132] indicate the audio hardware communication layer of the call device to load the sound data acquired by the recording device to an ADSP in the call device through a preset load key value pair, so that the ADSP extracts the first voiceprint feature of the sound data acquired by the recording device.
[0133] Optionally, the copying module 110 is configured to train a general background model according to the sound data collected by the recording device to obtain a trained voiceprint recognition model, and the first voiceprint feature of the sound data collected by the recording device;
[0134] The determining module 140 is configured to input the first voiceprint feature and the second voiceprint feature into the trained voiceprint recognition model to obtain a target voiceprint feature in the second voiceprint feature that matches the first voiceprint feature.
[0135] As to the apparatus in the above-mentioned embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the method, and will not be described in detail here.
[0136] In addition, it is worth noting that the modules in the above-mentioned embodiments can be independent devices or the same device when implemented, for example, the determining module 140 and the generating module 150 can be the same module or two modules, and the present disclosure does not limit this.
[0137] The present disclosure also provides a call device, comprising:
[0138] a processor;
[0139] a memory for storing processor-executable instructions;
[0140] wherein the processor is configured to:
[0141] copy the first voiceprint feature pre-stored by the application side of the call device to an audio digital signal processor module ADSP of the call device through shared memory before the call device starts a call;
[0142] in response to a noise reduction enabling operation on a call interface, initialize a speech algorithm in the ADSP during the call process of the call device, wherein the initialized speech algorithm can establish communication between the ADSP and the application side; and
[0143] transmit sound data in the call process from an audio hardware layer of the call device to the ADSP based on the communication, so that the ADSP extracts a second voiceprint feature of the sound data based on the initialized speech algorithm;
[0144] determine a target voiceprint feature that matches the first voiceprint feature from the second voiceprint feature based on the initialized speech algorithm in the ADSP;
[0145] generate speech call data according to the target voiceprint feature.
[0146] The present disclosure also provides a computer readable storage medium, having stored thereon computer program instructions, which when executed by a processor implement the steps of the voiceprint noise reduction method provided by the present disclosure.
[0147] Figure 16 is a block diagram of an apparatus 800 for voiceprint noise reduction according to an exemplary embodiment. The apparatus 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0148] Referring to Figure 16 The apparatus 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814 and a communication component 816.
[0149] The processing component 802 usually controls overall operations of the apparatus 800, such as operations associated with displaying, making phone calls, data communications, camera operations and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the voiceprint noise reduction method described above. Further, the processing component 802 can include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0150] The memory 804 is configured to store various types of data to support operations of the apparatus 800. Examples of these data include instructions for any applications or methods operating on the apparatus 800. The memory 804 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic storage devices, flash memory, magnetic disks or optical disks.
[0151] The power supply component 806 supplies electrical power for the various components of the apparatus 800. The power supply component 806 can include a power supply management system, one or more power sources, and other components associated with generating, managing and distributing power for the apparatus 800.
[0152] The multimedia component 808 includes a screen providing an output interface between the device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors for sensing a touch, a slide and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data.
[0153] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the device 800 is in an operation mode, such as a call mode, a recording mode and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting an audio signal.
[0154] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, such as a keypad, a click wheel, buttons and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button and a lock button.
[0155] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the device 800. For example, the sensor component 814 can detect an open / closed position of the device 800, relative positioning of components, such as a display and a keypad of the device 800, a change of position of the device 800 or a component of the device 800, presence or absence of user contact with the device 800, a change in orientation of the device 800 or acceleration / deceleration of the device 800, and a temperature change of the device 800, among a plethora of other examples. The sensor component 814 can include an orientation sensor, an acceleration sensor, a proximity sensor, a gesture sensor, a gravity sensor, a biometric sensor, a temperature sensor, a humidity sensor, and an illuminance sensor, among a plethora of other examples.
[0156] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate close proximity communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0157] In an exemplary embodiment, the device 800 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic components, for performing the voiceprint noise reduction method described above.
[0158] In an exemplary embodiment, a non-transitory computer readable storage medium, such as the memory 804 including instructions, is also provided, which can be executed by the processor 820 of the device 800 to complete the voiceprint noise reduction method described above. For example, the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0159] In another exemplary embodiment, a computer program product is also provided, which contains a computer program capable of being executed by a programmable device, and the computer program has code portions for executing the voiceprint noise reduction method described above when executed by the programmable device.
[0160] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure. It is intended that the present disclosure cover any and all variations of the present disclosure including those variations which can be incorporated into the above detailed description and making use of the general principles of the present disclosure. It is intended that the present disclosure include all such as fall within the scope of the appended claims and their equivalents. The specification and examples given are intended as illustrative only and not limiting of the true scope and spirit of the present disclosure.
[0161] It should be understood that the present disclosure is not limited to the precise structures described and shown in the drawings, and that various modifications and changes can be made to the embodiments without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the claims appended hereto.
Claims
1. A voiceprint noise reduction method, characterized in that, The method comprises the following steps: Before the call device starts a call, a first voiceprint feature pre-stored in an application side of the call device is copied into an audio digital signal processor module (ADSP) of the call device through a shared memory, wherein, after a static service receives a message command for indicating that the copying of the first voiceprint feature into the shared memory area is completed through a message processing function, the first voiceprint feature in the shared memory area is copied into a static service buffer of the ADSP; During the call process of the call device, a voice algorithm in the ADSP is initialized in response to a noise reduction enabling operation on a call interface, wherein, the initialized voice algorithm can establish a communication between the ADSP and the application side; and Sound data in the call process is transmitted from an audio hardware layer of the call device to the ADSP based on the communication, so that the ADSP extracts a second voiceprint feature of the sound data based on the initialized voice algorithm; Based on the initialized voice algorithm in the ADSP, a target voiceprint feature matching the first voiceprint feature is determined from the second voiceprint feature; Voice call data is generated according to the target voiceprint feature.
2. The method of claim 1, wherein, The first voiceprint feature is pre-stored in a kernel of the application side, and the copying of the first voiceprint feature pre-stored in the application side of the call device into the ADSP of the call device through the shared memory before the call device starts the call comprises the following steps: A shared memory area between the kernel and the ADSP is created; The first voiceprint feature pre-stored in the kernel is copied into the shared memory area; After the copying of the first voiceprint feature into the shared memory area is completed, the first voiceprint feature in the shared memory area is copied into a static service buffer of the ADSP.
3. The method of claim 2, wherein, The first voiceprint feature is stored in the static service buffer in a pre-defined storage manner, wherein, the pre-defined storage manner comprises storing the first voiceprint feature in a 32-bit compilation manner and a 64-bit compilation manner respectively.
4. The method according to any one of claims 1 to 3, characterized in that, The transmission of the sound data in the call process from the audio hardware layer of the call device to the ADSP based on the communication comprises the following steps: In a case where it is determined that the mode of the call is a network call mode, the sound data is acquired from an audio path of an audio hardware communication layer of the call device based on the application side, and the sound data is transmitted to the ADSP through an audio calibration communication between the application side and the ADSP.
5. The method according to any one of claims 1-3, characterized in that, In a case where it is determined that the mode of the call is a direct external dialing mode, a target calling interface of an audio hardware communication layer of the call device is determined from a debugging interface based on a calling log of a preset audio calibration tool calling debugging interface through the application side. obtain the sound data from the audio hardware communication layer according to the target calling interface; transmit the sound data to the ADSP through audio calibration communication between the application side and the ADSP.
6. The method according to any one of claims 1-3, characterized in that, The obtaining of the first voiceprint feature comprises: in response to a voiceprint recording operation on a setting interface of the call device, calling audio hardware through a preset calling key value pair, the calling key value pair being used to instruct the audio hardware to create a recording channel using a specified recording device; collecting sound data according to the recording device; loading the sound data collected by the recording device to an ADSP in the call device through a preset loading key value pair to enable the ADSP to extract the first voiceprint feature of the sound data collected by the recording device.
7. The method of claim 6, wherein, The ADSP extracting the first voiceprint feature of the sound data collected by the recording device comprises: training a general background model according to the sound data collected by the recording device to obtain a trained voiceprint recognition model and the first voiceprint feature of the sound data collected by the recording device; The determining of the target voiceprint feature matching the first voiceprint feature from the second voiceprint feature based on the initialized voice algorithm in the ADSP comprises: inputting the first voiceprint feature and the second voiceprint feature into the trained voiceprint recognition model to obtain the target voiceprint feature matching the first voiceprint feature from the second voiceprint feature.
8. A voiceprint noise reduction device, characterized in that, The device comprises: a copying module configured to copy a first voiceprint feature pre-stored by an application side of a call device to an audio digital signal processor module (ADSP) of the call device through shared memory before the call device starts a call, wherein a static service copies the first voiceprint feature in a shared memory area to a static service buffer area of the ADSP after receiving a message command representing completion of copying the first voiceprint feature to the shared memory area through a message processing function; an initialization module configured to initialize a voice algorithm in the ADSP in response to a noise reduction enabling operation on a call interface during a call process of the call device, wherein the initialized voice algorithm can establish communication between the ADSP and the application side; and a transmission module configured to transmit sound data in the call process from an audio hardware layer of the call device to the ADSP based on the communication to enable the ADSP to extract a second voiceprint feature of the sound data based on the initialized voice algorithm; a determination module configured to determine a target voiceprint feature matching the first voiceprint feature from the second voiceprint feature based on the initialized voice algorithm in the ADSP; a generation module configured to generate voice call data according to the target voiceprint feature.
9. A talker device, characterized by comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: Before the call device starts a call, a first voiceprint feature pre-stored on an application side of the call device is copied to an audio digital signal processor module (ADSP) of the call device through a shared memory, wherein, after a message command indicating that the copying of the first voiceprint feature to the shared memory area is completed is received by a static service through a message processing function, the first voiceprint feature in the shared memory area is copied to a static service buffer of the ADSP; During the call of the call device, in response to a noise reduction enabling operation on a call interface, a voice algorithm in the ADSP is initialized, wherein the initialized voice algorithm can establish communication between the ADSP and the application side; and Sound data during the call is transmitted from an audio hardware layer of the call device to the ADSP based on the communication, so that the ADSP extracts a second voiceprint feature of the sound data based on the initialized voice algorithm; Based on the initialized voice algorithm in the ADSP, a target voiceprint feature matching the first voiceprint feature is determined from the second voiceprint feature; Voice call data is generated according to the target voiceprint feature.
10. A computer-readable storage medium having stored thereon computer program instructions, wherein, The program instructions, when executed by a processor, implement the steps of the method of any one of claims 1-7.
Citation Information
Patent Citations
Conversation voice optimizing method, apparatus and conversation terminal
CN106920559A
Call method and device, storage medium and mobile terminal
CN107566658A
Network telephone voice equipment, processing method and device and storage medium
CN111200692A