Audio processing method and device
By obtaining the audio data and voiceprint feature information of the target object for voice detection and voiceprint analysis, the audio of non-target objects is eliminated, and the problem of voice leakage of users around them during audio calls is solved, improving user experience and privacy protection.
Patent Information
- Application Number
- CN202410076160.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-18
AI Technical Summary
During the audio call, the speech sounds of surrounding users are collected and upstreamed to the remote end, affecting the experience of other users and may disclose privacy. The existing methods rely on manual operations and pose a risk of privacy leakage.
By obtaining the current audio data of the target object and the target voiceprint feature information, performing voice detection and voiceprint analysis, determining whether the voiceprint feature of the target object is included, and if not included, audio processing is performed to eliminate audio that is not the target object.
Effectively avoid audio transmission of non-target objects, prevent privacy leakage, and improve user experience.
Smart Images

Figure CN120340544A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an audio processing method and apparatus. Background Art
[0002] Artificial intelligence is a technology that simulates human intelligence and aims to create intelligent machines that can autonomously learn and process information. The main directions of artificial intelligence software technology include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0003] With the development of the Internet, audio and video have become common application scenarios in daily life, such as audio calls and interactive live broadcasts. Real-time voice interaction can greatly facilitate people's information transmission and improve communication efficiency. In the actual use process, there will be such a scenario that when the target user is on a voice call, there are surrounding users nearby, and the target user does not speak or the target user needs to leave the acquisition device for a period of time due to some factors. If the microphone is forgotten to be turned off, the device will collect the voices of the surrounding users and upload them to the remote end, affecting the experience of other users and even leading to the leakage of the privacy of the surrounding users. In view of the above problems, taking the audio call scenario as an example, the existing methods mainly rely on the co-host to mute the users whose voices are collected and interfered; however, such operations are relatively cumbersome and require the host to operate repeatedly, and there is still a risk of privacy leakage if privacy is involved. Summary of the Invention
[0004] In view of the above existing technical problems, the present disclosure proposes an audio processing method and apparatus.
[0005] According to one aspect of the embodiments of the present disclosure, an audio processing method is provided, including:
[0006] Obtaining current audio data corresponding to a target object and target voiceprint feature information corresponding to the target object; the target voiceprint feature information is extracted from preset recorded audio data of the target object;
[0007] Performing voice detection on the current audio data to obtain a voice detection result corresponding to the current audio data;
[0008] In the case where the voice detection result indicates that the current audio data contains object audio features, performing voiceprint analysis on the current audio data based on the target voiceprint feature information to obtain a voiceprint analysis result corresponding to the current audio data;
[0009] In the case that the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, based on the first processing strategy, perform audio processing on the current audio data; the first processing strategy is used to indicate performing elimination processing on the current audio data.
[0010] According to another aspect of the embodiments of the present disclosure, there is provided an audio processing device, including:
[0011] A current data acquisition module, configured to acquire the current audio data corresponding to the target object and the target voiceprint feature information corresponding to the target object; the target voiceprint feature information is extracted from the preset recorded audio data of the target object;
[0012] A voice detection module, configured to perform voice detection on the current audio data to obtain a voice detection result corresponding to the current audio data;
[0013] A voiceprint analysis module, configured to, in the case that the voice detection result indicates that the current audio data includes object audio features, perform voiceprint analysis on the current audio data based on the target voiceprint feature information to obtain a voiceprint analysis result corresponding to the current audio data;
[0014] An audio processing module, configured to, in the case that the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, perform audio processing on the current audio data based on the first processing strategy; the first processing strategy is used to indicate performing elimination processing on the current audio data.
[0015] Optionally, the audio processing module includes:
[0016] An accumulated duration determination module, configured to, in the case that the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, determine an accumulated interference duration based on the audio time information corresponding to the current audio data and the current interference duration reset time;
[0017] A first audio processing module, configured to, in the case that the accumulated interference duration is greater than or equal to a preset duration, perform audio processing on the current audio data based on the first processing strategy.
[0018] Optionally, the accumulated duration determination module includes:
[0019] A first duration determination module, configured to, in the case that the current interference duration reset time is the start time corresponding to the current audio data, determine the accumulated interference duration based on the audio duration of the object audio features included in the current audio data.
[0020] Optionally, the cumulative duration determination module includes:
[0021] A historical duration acquisition module, configured to, when the current interference duration reset time is before the start time corresponding to the current audio data, acquire the historical interference duration corresponding to the first historical audio data; the first historical audio data is the audio data corresponding to the target object within a first time period, the first time period is the time period from the current interference duration reset time to the start time corresponding to the current audio data, and the historical interference duration is the audio duration including the object audio feature in the first historical audio data.
[0022] A second duration determination module, configured to determine the cumulative interference duration based on the historical interference duration and the audio duration including the object audio feature in the current audio data.
[0023] Optionally, the device further includes:
[0024] A second audio processing module, configured to, when the cumulative interference duration is less than the preset duration, perform audio processing on the current audio data based on a second processing strategy; the second processing strategy is used to indicate that no elimination processing is performed on the current audio data.
[0025] Optionally, the device further includes:
[0026] A third audio processing module, configured to, when the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, reset the cumulative interference duration and perform audio processing on the current audio data based on a second processing strategy.
[0027] Optionally, the device further includes:
[0028] A first noise reduction processing module, configured to perform noise reduction processing on the current audio data based on a preset noise template to obtain the noise-reduced current audio data.
[0029] Correspondingly, the voice detection module includes:
[0030] A detection result acquisition module, configured to perform voice detection on the noise-reduced current audio data to obtain the voice detection result.
[0031] Optionally, the third audio processing module includes:
[0032] A duration reset module, configured to, when the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, reset the cumulative interference duration.
[0033] The second noise reduction processing module is used to perform noise reduction processing on the current audio data after noise reduction based on the target voiceprint feature information to obtain the object audio data corresponding to the target object;
[0034] The fourth audio processing module is used to perform audio processing on the object audio data corresponding to the target object based on the second processing strategy.
[0035] Optionally, the target voiceprint feature information includes the first voiceprint feature information corresponding to multiple preset morphemes; the voiceprint analysis module includes:
[0036] A voice recognition module is used to perform voice recognition on the current audio data to obtain the current morpheme corresponding to the current audio data and the target audio data corresponding to the current morpheme; the target audio data is the valid audio data belonging to the current morpheme in the current audio data;
[0037] The voiceprint feature extraction module is used to perform voiceprint feature extraction processing on the target audio data when there is an intersection between the multiple preset morphemes and the current morpheme to obtain the second voiceprint feature information corresponding to the current morpheme;
[0038] The current feature determination module is used to determine the third voiceprint feature information corresponding to the current morpheme from the first voiceprint feature information corresponding to the multiple preset morphemes;
[0039] The voiceprint feature matching module is used to perform matching processing on the third voiceprint feature information and the second voiceprint feature information to obtain a target matching result;
[0040] The analysis result generation module is used to generate the voiceprint analysis result based on the target matching result.
[0041] Optionally, the target matching result includes a first matching result, and the first matching result is used to indicate that there is no match between the third voiceprint feature information and the second voiceprint feature information; the analysis result generation module includes:
[0042] The cumulative quantity determination module is used to determine the cumulative morpheme quantity based on the audio time information corresponding to the current audio data and the current matching morpheme reset time when the target matching result is the first matching result;
[0043] The first result generation module is used to determine that the voiceprint analysis result is the first analysis result when the cumulative morpheme quantity is greater than or equal to the preset morpheme quantity; the first analysis result is used to indicate that the current audio data does not include any voiceprint feature information of the target object.
[0044] Optionally, the cumulative quantity determination module includes:
[0045] A first quantity determination module, configured to, when the current matching morpheme reset time is the start time corresponding to the current audio data, determine the target morpheme quantity corresponding to the current audio data based on the current morpheme; the target morpheme quantity is the quantity of at least one morpheme included in the current morpheme and belonging to the multiple preset morphemes;
[0046] A second quantity determination module, configured to determine the cumulative morpheme quantity based on the target morpheme quantity.
[0047] Optionally, the cumulative quantity determination module includes:
[0048] A historical quantity acquisition module, configured to, when the current matching morpheme reset time is before the start time corresponding to the current audio data, acquire the historical morpheme quantity corresponding to second historical audio data; the second historical audio data is the audio data corresponding to the target object within a second time period, the second time period is the time period from the current matching morpheme reset time to the start time corresponding to the current audio data, the historical morpheme quantity is the quantity of at least one morpheme included in the historical morphemes corresponding to the second historical audio data and belonging to the multiple preset morphemes, and the historical morphemes corresponding to the second historical audio data are obtained by performing speech recognition on the second historical audio data;
[0049] A third quantity determination module, configured to determine the target morpheme quantity corresponding to the current audio data based on the current morpheme; the target morpheme quantity is the quantity of at least one morpheme included in the current morpheme and belonging to the multiple preset morphemes;
[0050] A fourth quantity determination module, configured to determine the cumulative morpheme quantity based on the historical morpheme quantity and the target morpheme quantity.
[0051] Optionally, the analysis result generation module includes:
[0052] A second result generation module, configured to, when there is no intersection between the multiple preset morphemes and the current morpheme, determine that the voiceprint analysis result is a second analysis result; the second analysis result is used to indicate that it is impossible to determine whether the voiceprint feature of the target object exists in the current audio data.
[0053] Optionally, the apparatus further includes:
[0054] A fifth audio processing module, configured to perform audio processing on the current audio data based on a current processing strategy when the voiceprint analysis result is the second analysis result; the current processing strategy is the first processing strategy or the second processing strategy.
[0055] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the above audio processing method.
[0056] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the above audio processing method.
[0057] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product containing instructions, when it runs on a computer, enabling the computer to execute the above audio processing method.
[0058] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0059] By obtaining the current audio data corresponding to the target object and the target voiceprint feature information corresponding to the target object, where the target voiceprint feature information is extracted from the preset recorded audio data of the target object, the current audio data of the target object and its voiceprint features can be obtained. Then, voice detection is performed on the current audio data to obtain a voice detection result corresponding to the current audio data, which can realize the detection of whether the current audio data contains object audio features. Next, when the voice detection result indicates that the current audio data contains object audio features, voiceprint analysis is performed on the current audio data based on the target voiceprint feature information to obtain a voiceprint analysis result corresponding to the current audio data, which can realize the analysis of whether the current audio data contains the voiceprint features of the target object. Then, when the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, audio processing is performed on the current audio data based on the first processing strategy, where the first processing strategy is used to indicate elimination processing of the current audio data, which can avoid the transmission of audio of non-target objects in the current audio data and the risk of privacy leakage of non-target objects, thereby improving the user experience.
[0060] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The accompanying drawings herein are incorporated into and form a part of the specification, showing embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure, and do not constitute an undue limitation of the present disclosure.
[0062] Figure 1 is a schematic diagram of an application system shown according to an exemplary embodiment;
[0063] Figure 2 is a flowchart of an audio processing method shown according to an exemplary embodiment;
[0064] Figure 3 is a schematic flow diagram of an audio processing method shown according to an exemplary embodiment;
[0065] Figure 4 is a block diagram of an audio processing device shown according to an exemplary embodiment;
[0066] Figure 5 is a block diagram of an electronic device for audio processing of current audio data shown according to an exemplary embodiment;
[0067] Figure 6 is a block diagram of another electronic device for audio processing of current audio data shown according to an exemplary embodiment. Detailed Embodiments
[0068] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0069] The special term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.
[0070] In addition, for a better description of the present application, numerous specific details are given in the following detailed embodiments. Those skilled in the art should understand that the present application can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present application.
[0071] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0072] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning in several major directions.
[0073] The key technologies of speech technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the most promising human-computer interaction methods in the future.
[0074] Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0075] As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtual machines containing operating systems), storage devices, and network devices.
[0076] According to the logical function division, on the IaaS (Infrastructure as a Service) layer, the PaaS (Platform as a Service) layer can be deployed. Above the PaaS layer, the SaaS (Software as a Service) layer can be deployed. Or the SaaS can be directly deployed on the IaaS. PaaS is the platform for software operation, such as databases, web containers, etc. SaaS are various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are the upper layers relative to IaaS.
[0077] Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through functions such as cluster applications, grid technology, and distributed storage file systems, and works together through application software or application interfaces to jointly provide data storage and business access functions to the outside world.
[0078] Currently, the storage method of the storage system is as follows: Create a logical volume. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may be composed of the disks of a certain storage device or several storage devices. The client stores data on a certain logical volume, that is, stores the data on the file system. The file system divides the data into many parts, and each part is an object. The object not only contains data but also additional information such as data identification (ID, ID entity). The file system writes each object into the physical storage space of the logical volume respectively, and the file system will record the storage location information of each object. Thus, when the client requests to access the data, the file system can enable the client to access the data according to the storage location information of each object.
[0079] The process of the storage system allocating physical storage space for a logical volume is specifically as follows: According to the capacity estimation of the objects stored in the logical volume (this estimation often has a large margin relative to the actual capacity of the objects to be stored) and the group of the redundant array of independent disks (RAID, Redundant Array of Independent Disk), the physical storage space is pre-divided into stripes. A logical volume can be understood as a stripe, thereby allocating physical storage space for the logical volume.
[0080] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0081] The solution provided in the embodiments of this application relates to technologies such as speech processing technology of artificial intelligence, and is specifically described through the following embodiments:
[0082] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application system shown according to an exemplary embodiment. The application system can be used for the audio processing method of this application. As Figure 1 shown, the application system can at least include a server 01 and a terminal 02.
[0083] In the embodiments of this application, the server 01 can be used to send the processed audio data to the terminal to be received. Specifically, the above-mentioned server 01 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0084] In the embodiments of this application, the terminal 02 can be used to perform audio processing on the current audio data. The above-mentioned terminal 02 can include entity devices such as smartphones, desktop computers, tablet computers, laptop computers, smart speakers, vehicle-mounted terminals, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, etc., or can also include software running on the entity devices, such as application programs, etc. In the embodiments of this application, the operating systems running on the above-mentioned terminal 02 can include, but are not limited to, Android system, IOS system, linux, windows, etc.
[0085] In addition, it should be noted that Figure 1 what is shown is only an application environment provided by this disclosure. In actual applications, there may also be other application environments. For example, the audio processing process of the current audio data can also be implemented on the server 01.
[0086] In the embodiments of this specification, the above-mentioned terminal 02 and the server 01 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any limitations in this regard.
[0087] It should be noted that the following figure shows a possible step sequence, and in fact, it is not necessarily limited to strictly follow this sequence. Some steps can be executed in parallel without depending on each other.
[0088] Specifically, Figure 2 is a flowchart of an audio processing method shown according to an exemplary embodiment. As Figure 2 shown, this audio processing method can be used in electronic devices such as terminals or servers, and specifically may include the following steps:
[0089] S201: Obtain the current audio data corresponding to the target object and the target voiceprint feature information corresponding to the target object.
[0090] In a specific embodiment, the target object may refer to an object for which audio needs to be collected. Specifically, taking the audio call scenario as an example, the voice of the target object can be collected through the audio collection module of the target terminal. Correspondingly, the target terminal can generate the voice data of the target object and send the above voice data of the target object to the server, so that the server sends the above voice data of the target object to the terminal of the call object; where the above call object may refer to an object that establishes a voice call with the target object.
[0091] In a specific embodiment, the current audio data corresponding to the target object may refer to the audio data currently collected by the target terminal. The current audio data may be audio data collected by the target terminal within the current time range. Among them, the target terminal can be used to collect the voice of the target object and upload it to the server. Exemplarily, the time length of the current time range can be 100 ms or 200 ms, etc. Specifically, the current audio data corresponding to the target object may not include the object voice feature, or may include the object voice feature; in the case where the current audio data includes the object voice feature, the current audio data corresponding to the target object may include the voiceprint feature of the target object, or may be the voiceprint feature of a non-target object. Among them, the non-target object may refer to an object that affects the voice collection of the target object; specifically, the non-target object may include objects located around the target object. It can be understood that during the voice collection process, in the presence of a non-target object, if the non-target object speaks, the voice of the non-target object may be collected, that is, the current voice data may include the voiceprint feature of the non-target object; or it may also be that while collecting the voice of the target object, the voice of the non-target object is collected, that is, the current audio data may include the voiceprint feature of the target object and may also include the voiceprint feature of the non-target object.
[0092] In a specific embodiment, when the target object enables the audio collection function, the current audio data can be collected through the audio collection module in the target terminal. Correspondingly, the target terminal can generate the current audio data.
[0093] In a specific embodiment, the target voiceprint feature information corresponding to the target object can be used to characterize the voiceprint feature of the target object. The target voiceprint feature information can include the first voiceprint feature information corresponding to each of a plurality of preset morphemes. The target voiceprint feature information can be extracted from the preset input audio data of the target object. The preset input audio data can be obtained by the target object through voice input based on the preset text information. Among them, the first voiceprint feature information corresponding to any preset morpheme can be used to characterize the voiceprint feature of the target object.
[0094] In a specific embodiment, before starting to collect the audio data of the target object, the preset input audio data of the target object can be obtained first; correspondingly, based on the preset text information, feature extraction processing can be performed on the preset input audio data to obtain the target voiceprint feature information. Among them, the preset text information can include a plurality of preset morphemes; the plurality of preset morphemes can include a plurality of preset common morphemes. Exemplarily, the plurality of preset morphemes can include common morphemes such as "I", "of", "already", and "you". Specifically, based on the preset text information, speech recognition can be performed on the preset input audio data to obtain the preset audio data corresponding to each of the plurality of preset morphemes in the preset text information; then, voiceprint feature extraction processing can be performed on the preset audio data corresponding to each preset morpheme to obtain the first voiceprint feature information corresponding to each of the above preset morphemes.
[0095] S203: Perform voice detection on the current audio data to obtain the voice detection result corresponding to the current audio data.
[0096] In a specific embodiment, the voice detection result can be used to indicate whether the current audio data of the target object contains object audio features. The voice detection result can include a first detection result and a second detection result. Specifically, the first detection result can be used to indicate that the current audio data contains object audio features. The second detection result can be used to indicate that the current audio data does not contain object audio features. Among them, the object audio features can be used to indicate the audio type included in the audio data; exemparily, the object audio features can include human voice features. Exemplarily, taking the voice call scenario as an example, when the voice detection result is the first detection result, it can be determined that the current audio data contains human voices; when the voice detection result is the second detection result, it can be determined that the current audio data does not contain human voices.
[0097] In a specific embodiment, data conversion processing can be performed on each frame of audio data in the above-mentioned current audio data to obtain frequency-domain data corresponding to each frame of audio data in the above-mentioned current audio data; then, based on a preset frequency range, energy ratio analysis can be performed on the frequency-domain data corresponding to each frame of audio data to obtain energy ratio index data corresponding to each frame of audio data; by comparing the energy ratio index data corresponding to each frame of audio data with preset ratio index data, a voice detection result corresponding to the current audio data can be obtained. Among them, the preset frequency range can refer to the frequency range of the target voice; for example, the preset frequency range can be a frequency less than or equal to 2KHZ. The energy ratio index data corresponding to any frame of audio data can represent the ratio between the energy belonging to the preset frequency range in the above-mentioned any frame of audio data and the total energy corresponding to the above-mentioned any frame of audio data. Specifically, in the case where the energy ratio index data of any frame of audio data is greater than or equal to the preset ratio index data, the voice detection result can be determined as the first detection result; in the case where the energy ratio index data of each frame of the above-mentioned audio data is less than the preset ratio index data, the voice detection result can be determined as the second detection result.
[0098] In a specific embodiment, the above method may further include:
[0099] Perform noise reduction processing on the current audio data based on a preset noise template to obtain the current audio data after noise reduction;
[0100] Correspondingly, the above-mentioned voice detection of the current audio data to obtain the voice detection result corresponding to the current audio data includes:
[0101] Perform voice detection on the current audio data after noise reduction to obtain the voice detection result.
[0102] In a specific embodiment, the preset noise template can be used to provide a reference for the noise reduction processing of the current audio data. The preset noise template can include one or more of noise templates such as a preset wind sound template, a preset rain sound template, a preset keyboard sound template, and a preset object noise template.
[0103] In a specific embodiment, the current audio data after noise reduction can refer to the current audio data after noise reduction processing.
[0104] In a specific embodiment, noise matching can be performed on the current audio data and a preset noise template to obtain a noise matching result. Correspondingly, based on the noise matching result, the matched noise in the current audio data can be eliminated to obtain the current audio data after noise reduction. Specifically, the preset noise template may or may not include a preset object noise template. It can be understood that in the case where the preset noise template includes a preset object noise template, the object noise (such as human voice noise) in the current audio data can be eliminated, that is, through noise reduction processing, the audio belonging to non-target objects in the current audio data can be eliminated, that is, the current audio data after noise reduction may not include the voiceprint features of non-target objects; in the case where the preset noise template does not include a preset object noise template and the current audio data includes the voiceprint features of non-target objects, the current audio data after noise reduction will still include the voiceprint features of the above non-target objects.
[0105] In a specific embodiment, the current audio data after noise reduction can be subjected to echo cancellation processing, and the audio data after echo cancellation processing can be subjected to audio gain processing to obtain the current audio data after preprocessing; correspondingly, the voice detection can be performed on the current audio data after the above preprocessing to obtain the above voice detection result.
[0106] In the above embodiment, based on the preset noise template, the current audio data is subjected to noise reduction processing to obtain the current audio data after noise reduction, and the voice detection is performed on the current audio data after noise reduction to obtain the voice detection result, which can reduce the interference of the noise in the current audio data on the subsequent voiceprint analysis.
[0107] S205: When the voice detection result indicates that the current audio data contains object audio features, based on the target voiceprint feature information, voiceprint analysis is performed on the current audio data to obtain the voiceprint analysis result corresponding to the current audio data.
[0108] In a specific embodiment, the voiceprint analysis result corresponding to the current audio data can be used to indicate whether the current audio data includes the voiceprint feature information of the target object. The voiceprint analysis result corresponding to the current audio data may include a first analysis result, a second analysis result, or a third analysis result. Among them, the first analysis result can be used to indicate that the current audio data does not include any voiceprint feature information of the target object. The second analysis result can be used to indicate that it is impossible to determine whether there is voiceprint feature information of the target object in the current audio data. The third analysis result can be used to indicate that the current audio data includes any voiceprint feature information of the target object.
[0109] In a specific embodiment, the above step S205 may include:
[0110] Perform speech recognition on the current audio data to obtain the current morpheme corresponding to the current audio data and the target audio data corresponding to the current morpheme;
[0111] In the case where there is an intersection between multiple preset morphemes and the current morpheme, perform voiceprint feature extraction processing on the target audio data to obtain the second voiceprint feature information corresponding to the current morpheme;
[0112] Determine the third voiceprint feature information corresponding to the current morpheme from the first voiceprint feature information corresponding to each of the multiple preset morphemes;
[0113] Perform matching processing on the third voiceprint feature information and the second voiceprint feature information to obtain the target matching result;
[0114] Generate a voiceprint analysis result based on the target matching result.
[0115] In a specific embodiment, the current morpheme may refer to at least one morpheme included in the speech text information corresponding to the current audio data. Specifically, the current morpheme may include at least one morpheme. It can be understood that the morphemes in the current morpheme may be the speech of the target object or the speech of the non-target object.
[0116] In a specific embodiment, the target audio data may refer to the valid audio data in the current audio data that belongs to the current morpheme. The target audio data may include the valid audio data corresponding to at least one morpheme in the current morpheme.
[0117] In a specific embodiment, the current audio data can be segmented to obtain multiple audio segments to be recognized. Any audio segment to be recognized can be an audio segment that may contain the complete waveform belonging to a single morpheme; then, perform morpheme recognition processing on each audio segment to be recognized to obtain the above-mentioned current morpheme; then, perform valid audio extraction on the audio segments to be recognized corresponding to each morpheme in the current morpheme to obtain the valid audio data corresponding to at least one morpheme in the above-mentioned current morpheme; correspondingly, the above-mentioned target audio data can be obtained.
[0118] In a specific embodiment, when the number of morphemes included in the current morpheme is one, the first comparison result can be obtained by comparing the current morpheme with each of the multiple preset morphemes; in the case where the first comparison result indicates that the current morpheme is the same as any preset morpheme, it can be determined that there is an intersection between the multiple preset morphemes and the current morpheme; in the case where the first comparison result indicates that the current morpheme is different from each preset morpheme, it can be determined that there is no intersection between the above-mentioned current morpheme and the multiple preset morphemes.
[0119] In a specific embodiment, when the number of morphemes included in the current morpheme is multiple, the second comparison result can be obtained by sequentially comparing each morpheme in the current morpheme with each morpheme in multiple preset morphemes; when the second comparison result indicates that there is a morpheme in the current morpheme that is the same as any one of the preset morphemes, it can be determined that there is an intersection between the multiple preset morphemes and the current morpheme; when the second comparison result indicates that there is no morpheme in the current morpheme that is the same as any one of the preset morphemes, it can be determined that there is no intersection between the current morpheme and the multiple preset morphemes.
[0120] In a specific embodiment, the second voiceprint feature information can be used to characterize the voiceprint feature of the target audio data corresponding to the current morpheme. The second voiceprint feature information can include the voiceprint feature information corresponding to at least one morpheme in the current morpheme. It can be understood that the voiceprint feature information of different objects is different, and based on the above first voiceprint feature information and second voiceprint feature information, it can be determined whether the voiceprint feature of the target object is included in the current audio data. Among them, any voiceprint feature information can include one or more of frequency distribution index data and loudness difference index data, etc. The frequency distribution index data can characterize the energy distribution of the corresponding audio data at different frequencies; the loudness difference index data can characterize the loudness difference of the corresponding audio data at different frequencies. Exemplarily, any frequency distribution index data can include energy data corresponding to frequency f1, energy data corresponding to frequency f2, energy data corresponding to frequency f3, etc., which are the energy data corresponding to multiple different frequencies.
[0121] In a specific embodiment, when there is an intersection between the multiple preset morphemes and the current morpheme, data conversion processing can be performed on the effective audio data corresponding to each morpheme in the target audio data to obtain the morpheme audio frequency domain data corresponding to each morpheme in the target audio data; frequency distribution analysis can be performed on the morpheme audio frequency domain data corresponding to each morpheme in the target audio data to obtain the frequency distribution index data corresponding to each morpheme in the target audio data; loudness difference analysis can be performed on the morpheme audio frequency domain data corresponding to each morpheme in the target audio data to obtain the loudness difference index data corresponding to each morpheme in the target audio data.
[0122] In a specific embodiment, the third voiceprint feature information can be used to provide a reference for the matching of the second voiceprint feature information corresponding to the current morpheme to achieve voiceprint analysis.
[0123] In a specific embodiment, when the number of morphemes included in the current morpheme is one, the voiceprint feature information corresponding to the current morpheme in the first voiceprint feature information corresponding to each of the multiple preset morphemes can be used as the third voiceprint feature information.
[0124] In a specific embodiment, when the number of morphemes included in the current morpheme is multiple, the current intersection morpheme in the current morpheme can be determined based on multiple preset morphemes and the current morpheme first; the first acoustic feature information corresponding to the current intersection morpheme can be selected from the first acoustic feature information corresponding to each of the multiple preset morphemes as the third acoustic feature information corresponding to the current morpheme. Among them, the current intersection morpheme can refer to the morphemes in the current morpheme that belong to the multiple preset morphemes. The current intersection morpheme can include multiple morphemes.
[0125] In a specific embodiment, the target matching result can be used to indicate whether there is a match between the third acoustic feature information and the second acoustic feature information. The target matching result can include a first matching result or a second matching result. The first matching result can be used to indicate that there is no match between the third acoustic feature information and the second acoustic feature information. The second matching result can be used to indicate that there is a matching acoustic feature information between the third acoustic feature information and the second acoustic feature information.
[0126] In a specific embodiment, the target matching result can be obtained by performing a matching process on the acoustic feature information corresponding to the same morphemes in the third acoustic feature information and the second acoustic feature information.
[0127] In a specific embodiment, when the target matching result is the first matching result, the acoustic analysis result can be determined as the first analysis result.
[0128] In a specific embodiment, generating the acoustic analysis result based on the target matching result can include:
[0129] When the target matching result is the first matching result, determine the cumulative number of morphemes based on the audio time information corresponding to the current audio data and the current matching morpheme reset time;
[0130] When the cumulative number of morphemes is greater than or equal to the preset number of morphemes, determine the acoustic analysis result as the first analysis result.
[0131] In a specific embodiment, the audio time information corresponding to the current audio data can characterize the time attribute of the current audio data. The audio time information corresponding to the current audio data may include the start time of the current audio data; further, the audio time information corresponding to the current audio data may further include the end time of the current audio data and the audio duration of the current audio data. Among them, the start time of the current audio data may refer to the time corresponding to the first frame of audio data in the current audio data. The end time of the current audio data may refer to the time corresponding to the last frame of audio data in the current audio data. The audio duration of the current audio data may refer to the duration between the time of the first frame of audio data and the time of the last frame of audio data of the current audio data.
[0132] In a specific embodiment, the matching morpheme reset time may refer to the time when the reset operation of the cumulative morpheme count is performed. The current matching morpheme reset time may refer to the matching morpheme reset time closest to the start time of the current audio data.
[0133] In a specific embodiment, the cumulative morpheme count may refer to the number of unmatched morphemes accumulated from the current matching morpheme reset time to the end time of the current audio data.
[0134] In a specific embodiment, determining the cumulative morpheme count based on the audio time information corresponding to the current audio data and the current matching morpheme reset time may include:
[0135] In the case where the current matching morpheme reset time is the start time corresponding to the current audio data, determining the target morpheme count corresponding to the current audio data based on the current morpheme;
[0136] Determining the cumulative morpheme count based on the target morpheme count.
[0137] In a specific embodiment, the target morpheme count may refer to the number of at least one morpheme in the current morpheme that is included in multiple preset morphemes. Specifically, the current intersection morpheme may be determined based on the current morpheme and multiple preset morphemes; the target morpheme count may be determined based on the number of the current intersection morphemes.
[0138] In a specific embodiment, the target morpheme count may be used as the cumulative morpheme count.
[0139] In a specific embodiment, determining the cumulative morpheme count based on the audio time information corresponding to the current audio data and the current matching morpheme reset time may include:
[0140] In the case where the current matching morpheme reset time is before the start time corresponding to the current audio data, obtaining the historical morpheme count corresponding to the second historical audio data;
[0141] Determine the number of target morphemes corresponding to the current audio data based on the current morpheme.
[0142] Determine the cumulative number of morphemes based on the historical number of morphemes and the number of target morphemes.
[0143] In a specific embodiment, the second historical audio data may refer to the audio data corresponding to the target object within a second time period. Wherein, the second time period may refer to the time period between the current matching morpheme reset time and the start time corresponding to the current audio data.
[0144] In a specific embodiment, the historical number of morphemes may refer to the number of at least one morpheme included in the historical morphemes corresponding to the second historical audio data among multiple preset morphemes. Wherein, the historical morphemes corresponding to the second historical audio data are obtained by performing speech recognition on the second historical audio data. Exemplarily, when the speech text information corresponding to the second historical audio data is "I heard that the weather is nice tomorrow", the historical morphemes corresponding to the above second historical audio data may include "I", "heard", "that", "tomorrow", "the", "weather", "is", "nice"; assuming that the multiple preset morphemes include "I" and "the", it can be determined that the number of historical morphemes corresponding to the above second historical audio data is 2.
[0145] In a specific embodiment, the second time period may be determined based on the start time of the current audio data and the current matching morpheme reset time; the second historical audio data corresponding to the target object may be obtained based on the above second time period; performing speech recognition on the above second historical audio data may obtain the historical morphemes corresponding to the second historical audio data; the number of historical morphemes corresponding to the second historical audio data may be determined based on the historical morphemes corresponding to the above second historical audio data and multiple preset morphemes.
[0146] In a specific embodiment, the number of target morphemes may refer to the number of at least one morpheme included in the current morpheme among multiple preset morphemes. Specifically, the current intersection morphemes may be determined based on the current morpheme and multiple preset morphemes; the number of target morphemes may be determined based on the number of the current intersection morphemes.
[0147] In a specific embodiment, the historical number of morphemes and the number of target morphemes may be superimposed to obtain the cumulative number of morphemes.
[0148] In a specific embodiment, the number of preset morphemes may be set according to actual application needs, and the present disclosure does not make a limitation.
[0149] In the above embodiments, when the target matching result is the first matching result, the cumulative morpheme quantity is determined based on the audio time information corresponding to the current audio data and the current matching morpheme reset moment. When the cumulative morpheme quantity is greater than or equal to the preset morpheme quantity, the voiceprint analysis result is determined to be the first analysis result, which can improve the accuracy of the voiceprint analysis of non-target objects, thereby avoiding voice interruptions caused by voiceprint detection errors and enhancing the user experience.
[0150] In a specific embodiment, when the cumulative morpheme quantity is less than the preset morpheme quantity, the current audio data can be processed based on the current processing strategy. The current processing strategy can be the first processing strategy or the second processing strategy.
[0151] In a specific embodiment, the above-mentioned voiceprint analysis of the current audio data based on the target voiceprint feature information to obtain the voiceprint analysis result corresponding to the current audio data may further include:
[0152] When there is no intersection between multiple preset morphemes and the current morpheme, the voiceprint analysis result is determined to be the second analysis result.
[0153] In a specific embodiment, assuming that the multiple preset morphemes include "I" and "of", and the current morpheme is "listen", it can be determined that there is no intersection between the multiple preset morphemes and the current morpheme; correspondingly, when there is no intersection between the multiple preset morphemes and the current morpheme, the voiceprint analysis result can be determined to be the second analysis result.
[0154] In a specific embodiment, when the target matching result is the second matching result, the voiceprint analysis result can be determined to be the third analysis result.
[0155] In a specific embodiment, when the voice detection result indicates that the current audio data does not contain the object audio feature, the current audio data can be processed based on the second processing strategy. The second processing strategy can be used to indicate that the current audio data is not subjected to elimination processing. Specifically, when the voice detection result indicates that the current audio data does not contain the object audio feature, the processing strategy corresponding to the current audio data can be determined to be the second processing strategy, and then, a preset audio processing operation can be performed on the current audio data. The preset audio processing operation can include operations such as mixing processing operation and encoding operation.
[0156] S207: When the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, the current audio data is processed based on the first processing strategy.
[0157] In a specific embodiment, the first processing strategy can be used to indicate that the current audio data is to be processed for elimination.
[0158] In a specific embodiment, when the voiceprint analysis result is the first analysis result, it can be determined that the processing strategy corresponding to the current audio data is the first processing strategy; correspondingly, the current audio data can be processed for elimination. Specifically, the above elimination processing can include various elimination processing methods such as replacing the current audio data with preset silent audio data or deleting the current audio data, so as to eliminate the object audio features of non-target objects included in the current audio data. Among them, the above preset silent audio data can be silent audio data with the same duration as the above current audio data. It can be understood that through the elimination processing (such as replacing the current audio data with preset silent audio data), other users can be prevented from hearing the audio content of non-target objects.
[0159] In a specific embodiment, after the elimination processing, a preset audio processing operation can be performed on the current audio data after elimination. Specifically, the current audio data after elimination can be encoded and the encoded audio data can be sent to the server, so that other terminals can download the current audio data after elimination from the server and perform decoding and playing; or it can also be that the current audio data after elimination is subjected to mixing processing, and then the mixed audio data is encoded, and then the encoded audio data is sent to the server, so that other terminals can download the mixed audio data from the server and perform decoding and playing.
[0160] In a specific embodiment, the above step S207 can include:
[0161] S2071: When the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, based on the audio time information corresponding to the current audio data and the current interference duration reset time, determine the cumulative interference duration;
[0162] S2072: When the cumulative interference duration is greater than or equal to the preset duration, based on the first processing strategy, perform audio processing on the current audio data.
[0163] In a specific embodiment, the interference duration reset time can refer to the time when the reset operation of the cumulative interference duration is performed. The current interference duration reset time can refer to the interference duration reset time closest to the start time of the current audio data.
[0164] In a specific embodiment, the cumulative interference duration can refer to the effective audio duration of the mismatched morphemes accumulated from the current interference duration reset time to the end time of the current audio data.
[0165] In a specific embodiment, determining the cumulative interference duration based on the audio time information corresponding to the current audio data and the current interference duration reset moment may include:
[0166] When the current interference duration reset moment is the starting moment corresponding to the current audio data, determining the cumulative interference duration based on the audio duration of the audio data including the target audio feature in the current audio data.
[0167] In a specific embodiment, when the current interference duration reset moment is the starting moment corresponding to the current audio data, the cumulative interference duration may be determined based on the audio duration of the valid audio data corresponding to the current intersection morpheme. Specifically, when the current interference duration reset moment is the starting moment corresponding to the current audio data, the audio duration of the valid audio data corresponding to the current intersection morpheme may be used as the cumulative interference duration.
[0168] In a specific embodiment, determining the cumulative interference duration based on the audio time information corresponding to the current audio data and the current interference duration reset moment may include:
[0169] When the current interference duration reset moment is before the starting moment corresponding to the current audio data, obtaining the historical interference duration corresponding to the first historical audio data;
[0170] Determining the cumulative interference duration based on the historical interference duration and the audio duration of the audio data including the target audio feature in the current audio data.
[0171] In a specific embodiment, the first historical audio data may be the audio data corresponding to the target object within the first time period. Wherein, the first time period may be the time period from the current interference duration reset moment to the starting moment corresponding to the current audio data.
[0172] In a specific embodiment, the historical interference duration may be the audio duration of the audio data including the target audio feature in the first historical audio data.
[0173] In a specific embodiment, the first time period may be determined based on the starting moment of the current audio data and the current interference duration reset moment; the first historical audio data corresponding to the target object may be obtained based on the first time period; performing speech recognition on the first historical audio data may obtain a plurality of first morphemes corresponding to the first historical audio data and the valid audio duration corresponding to each first morpheme; based on a plurality of preset morphemes, at least one second morpheme may be determined from the plurality of first morphemes, wherein any one of the second morphemes is a morpheme belonging to the plurality of preset morphemes among the plurality of first morphemes; correspondingly, the valid audio durations corresponding to at least one second morpheme may be superimposed to obtain the historical interference duration.
[0174] In a specific embodiment, the historical interference duration and the audio duration of the effective audio data corresponding to the current intersection morpheme in the current audio data may be superimposed to obtain the cumulative interference duration.
[0175] In a specific embodiment, when the cumulative interference duration is greater than or equal to a preset duration, it may be determined that the processing strategy corresponding to the current audio data is the first processing strategy; correspondingly, the current audio data may be subjected to elimination processing, and a preset audio processing operation may be performed on the current audio data after the elimination processing. Specifically, the preset duration may be set according to actual application needs, and the present disclosure does not make a limitation.
[0176] In a specific embodiment, the above method may further include:
[0177] S2073: When the cumulative interference duration is less than the preset duration, based on the second processing strategy, audio processing is performed on the current audio data.
[0178] In a specific embodiment, when the cumulative interference duration is less than the preset duration, it may be determined that the processing strategy corresponding to the current audio data is the second processing strategy; correspondingly, a preset audio processing operation may be performed on the current audio data.
[0179] In the above embodiment, when the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, based on the audio time information corresponding to the current audio data and the current interference duration reset time, the cumulative interference duration is determined. When the cumulative interference duration is greater than or equal to the preset duration, based on the first processing strategy, audio processing is performed on the current audio data, which can avoid the situation of interrupted sound caused by the error of the voiceprint analysis result, thereby improving the user experience.
[0180] In a specific embodiment, the above method may further include:
[0181] S208: When the voiceprint analysis result indicates that the current audio data includes any voiceprint feature information of the target object, the cumulative interference duration is reset, and based on the second processing strategy, audio processing is performed on the current audio data.
[0182] In a specific embodiment, when the voiceprint analysis result indicates that the current audio data includes any voiceprint feature information of the target object, the cumulative interference duration may be reset, and based on the second processing strategy, audio processing is performed on the current audio data. Specifically, the cumulative interference duration may be reset to 0, and when it is determined that the current audio data belongs to the second processing strategy, a preset audio processing operation is performed on the current audio data.
[0183] In a specific embodiment, when the voiceprint analysis result is the third analysis result, an operation of resetting the cumulative interference duration and the cumulative morpheme quantity may be performed, and based on the second processing strategy, audio processing is performed on the current audio data.
[0184] In a specific embodiment, when the preset noise template does not include the preset object noise template, and when the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, resetting the cumulative interference duration, and based on the second processing strategy, performing audio processing on the current audio data may include:
[0185] When the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, reset the cumulative interference duration;
[0186] Based on the target voiceprint feature information, perform noise reduction processing on the noise-reduced current audio data to obtain the object audio data corresponding to the target object;
[0187] Based on the second processing strategy, perform audio processing on the object audio data corresponding to the target object.
[0188] In a specific embodiment, the object audio data corresponding to the target object may refer to the audio data of the voice belonging to the target object in the current audio data.
[0189] In a specific embodiment, when the preset noise template does not include the preset object noise template, further noise reduction may be performed on the above-mentioned noise-reduced current audio data based on the target voiceprint feature information to eliminate the voiceprint features in the above-mentioned noise-reduced current audio data that do not belong to the target object, so as to obtain the object audio data corresponding to the target object; correspondingly, a preset audio processing operation may be performed on the above-mentioned object audio data corresponding to the target object.
[0190] In the above embodiment, by resetting the cumulative interference duration when the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, performing noise reduction processing on the noise-reduced current audio data based on the target voiceprint feature information to obtain the object audio data corresponding to the target object, and performing audio processing on the object audio data corresponding to the target object based on the second processing strategy, further noise reduction of the current audio data can be achieved, the signal-to-noise ratio of the current audio data can be improved, and thus the user experience can be further improved.
[0191] In a specific embodiment, the above method may further include:
[0192] S209: When the voiceprint analysis result is the second analysis result, perform audio processing on the current audio data based on the current processing strategy.
[0193] In a specific embodiment, the current processing strategy may be the first processing strategy or the second processing strategy.
[0194] In a specific embodiment, when the voiceprint analysis result is the second analysis result, the processing strategy at the previous moment of the current moment can be obtained, and the processing strategy at the previous moment of the current moment can be used as the current processing strategy. Correspondingly, based on the current processing strategy, audio processing can be performed on the current audio data. It can be understood that when the voiceprint analysis result is the second analysis result, the processing strategy of the previous moment can be maintained.
[0195] In the above embodiment, by obtaining the current audio data corresponding to the target object and the target voiceprint feature information corresponding to the target object, where the target voiceprint feature information is extracted from the preset recorded audio data of the target object, the acquisition of the current audio data of the target object and its voiceprint features can be realized. Then, voice detection is performed on the current audio data to obtain the voice detection result corresponding to the current audio data, and the detection of whether the current audio data contains object audio features can be realized. Next, when the voice detection result indicates that the current audio data contains object audio features, voiceprint analysis is performed on the current audio data based on the target voiceprint feature information to obtain the voiceprint analysis result corresponding to the current audio data, and the analysis of whether the current audio data contains the voiceprint features of the target object can be realized. Then, when the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, audio processing is performed on the current audio data based on the first processing strategy, where the first processing strategy is used to indicate elimination processing of the current audio data, which can avoid the transmission of the audio of non-target objects in the current audio data and the risk of privacy leakage of non-target objects, thereby improving the user experience.
[0196] Figure 3 is a flowchart of an audio processing method shown according to an exemplary embodiment. Specifically, as Figure 3As shown, obtain the current audio data corresponding to the target object and the target voiceprint feature information corresponding to the target object; the target voiceprint feature information is extracted from the preset input audio data of the target object; perform voice detection on the current audio data to obtain the voice detection result corresponding to the current audio data; in the case where the voice detection result indicates that the current audio data contains object audio features, based on the target voiceprint feature information, perform voiceprint analysis on the current audio data to obtain the voiceprint analysis result corresponding to the current audio data; in the case where the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, based on the audio time information corresponding to the current audio data and the current interference duration reset time, determine the cumulative interference duration; in the case where the cumulative interference duration is greater than or equal to the preset duration, based on the first processing strategy, perform audio processing on the current audio data; in the case where the cumulative interference duration is less than the preset duration, based on the second processing strategy, perform audio processing on the current audio data; the second processing strategy is used to indicate that no elimination processing is performed on the current audio data; in the case where the voiceprint analysis result indicates that the current audio data includes any voiceprint feature information of the target object, reset the cumulative interference duration, and based on the second processing strategy, perform audio processing on the current audio data; in the case where the voiceprint analysis result is the second analysis result, based on the current processing strategy, perform audio processing on the current audio data.
[0197] Figure 4 is a block diagram of an audio processing device shown according to an exemplary embodiment. As Figure 4 shown, the device may include:
[0198] A current data acquisition module 410, which can be used to acquire the current audio data corresponding to the target object and the target voiceprint feature information corresponding to the target object; the target voiceprint feature information is extracted from the preset input audio data of the target object;
[0199] A voice detection module 420, which can be used to perform voice detection on the current audio data to obtain the voice detection result corresponding to the current audio data;
[0200] A voiceprint analysis module 430, which can be used to, in the case where the voice detection result indicates that the current audio data contains object audio features, based on the target voiceprint feature information, perform voiceprint analysis on the current audio data to obtain the voiceprint analysis result corresponding to the current audio data;
[0201] An audio processing module 440, which can be used to, in the case where the voiceprint analysis result indicates that the current audio data does not include any voiceprint feature information of the target object, based on the first processing strategy, perform audio processing on the current audio data; the first processing strategy is used to indicate that elimination processing is performed on the current audio data.
[0202] In a specific embodiment, the above audio processing module 440 may include:
[0203] An accumulated duration determination module, which can be used to determine the accumulated interference duration based on the audio time information corresponding to the current audio data and the current interference duration reset moment when the voiceprint analysis result indicates that any voiceprint feature information of the target object is not included in the current audio data;
[0204] A first audio processing module, which can be used to perform audio processing on the current audio data based on a first processing strategy when the accumulated interference duration is greater than or equal to a preset duration.
[0205] In a specific embodiment, the above accumulated duration determination module may include:
[0206] A first duration determination module, which can be used to determine the accumulated interference duration based on the audio duration including the object audio feature in the current audio data when the current interference duration reset moment is the starting moment corresponding to the current audio data.
[0207] In a specific embodiment, the above accumulated duration determination module may include:
[0208] A historical duration acquisition module, which can be used to acquire the historical interference duration corresponding to the first historical audio data when the current interference duration reset moment is before the starting moment corresponding to the current audio data; the first historical audio data is the audio data corresponding to the target object within a first time period, the first time period is the time period from the current interference duration reset moment to the starting moment corresponding to the current audio data, and the historical interference duration is the audio duration including the object audio feature in the first historical audio data;
[0209] A second duration determination module, which can be used to determine the accumulated interference duration based on the historical interference duration and the audio duration including the object audio feature in the current audio data.
[0210] In a specific embodiment, the above device may further include:
[0211] A second audio processing module, which can be used to perform audio processing on the current audio data based on a second processing strategy when the accumulated interference duration is less than the preset duration; the second processing strategy is used to indicate that the current audio data is not subjected to elimination processing.
[0212] In a specific embodiment, the above device may further include:
[0213] The third audio processing module can be used to reset the accumulated interference duration when the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, and perform audio processing on the current audio data based on the second processing strategy.
[0214] In a specific embodiment, the above device may further include:
[0215] The first noise reduction processing module can be used to perform noise reduction processing on the current audio data based on a preset noise template to obtain the current audio data after noise reduction;
[0216] Correspondingly, the above voice detection module 420 may include:
[0217] The detection result acquisition module can be used to perform voice detection on the current audio data after noise reduction to obtain a voice detection result.
[0218] In a specific embodiment, the above third audio processing module may include:
[0219] The duration reset module can be used to reset the accumulated interference duration when the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data;
[0220] The second noise reduction processing module can be used to perform noise reduction processing on the current audio data after noise reduction based on the target voiceprint feature information to obtain the object audio data corresponding to the target object;
[0221] The fourth audio processing module can be used to perform audio processing on the object audio data corresponding to the target object based on the second processing strategy.
[0222] In a specific embodiment, the above voiceprint analysis module 430 may include:
[0223] The voice recognition module can be used to perform voice recognition on the current audio data to obtain the current morpheme corresponding to the current audio data and the target audio data corresponding to the current morpheme; the target audio data is the valid audio data in the current audio data that belongs to the current morpheme;
[0224] The voiceprint feature extraction module can be used to perform voiceprint feature extraction processing on the target audio data when there is an intersection between multiple preset morphemes and the current morpheme to obtain the second voiceprint feature information corresponding to the current morpheme;
[0225] The current feature determination module can be used to determine the third voiceprint feature information corresponding to the current morpheme from the first voiceprint feature information corresponding to each of the multiple preset morphemes;
[0226] The voiceprint feature matching module can be used to match the third voiceprint feature information and the second voiceprint feature information to obtain a target matching result;
[0227] The analysis result generation module can be used to generate a voiceprint analysis result based on the target matching result.
[0228] In a specific embodiment, the target matching result includes a first matching result, and the first matching result is used to indicate that there is no match between the third voiceprint feature information and the second voiceprint feature information; the above analysis result generation module may include:
[0229] The cumulative quantity determination module can be used to determine the cumulative morpheme quantity based on the audio time information corresponding to the current audio data and the current matching morpheme reset moment when the target matching result is the first matching result;
[0230] The first result generation module can be used to determine that the voiceprint analysis result is the first analysis result when the cumulative morpheme quantity is greater than or equal to the preset morpheme quantity; the first analysis result is used to indicate that any voiceprint feature information of the target object is not included in the current audio data.
[0231] In a specific embodiment, the above cumulative quantity determination module may include:
[0232] The first quantity determination module can be used to determine the target morpheme quantity corresponding to the current audio data based on the current morpheme when the current matching morpheme reset moment is the starting moment corresponding to the current audio data; the target morpheme quantity is the quantity of at least one morpheme included in the current morpheme among multiple preset morphemes;
[0233] The second quantity determination module can be used to determine the cumulative morpheme quantity based on the target morpheme quantity.
[0234] In a specific embodiment, the above cumulative quantity determination module may include:
[0235] The historical quantity acquisition module can be used to acquire the historical morpheme quantity corresponding to the second historical audio data when the current matching morpheme reset moment is before the starting moment corresponding to the current audio data; the second historical audio data is the audio data corresponding to the target object within the second time period, the second time period is the time period from the current matching morpheme reset moment to the starting moment corresponding to the current audio data, the historical morpheme quantity is the quantity of at least one morpheme included in the historical morphemes corresponding to the second historical audio data among multiple preset morphemes, and the historical morphemes corresponding to the second historical audio data are obtained by performing speech recognition on the second historical audio data;
[0236] A third quantity determination module, which can be used to determine the number of target morphemes corresponding to the current audio data based on the current morpheme; the number of target morphemes is the number of at least one morpheme included in the current morpheme among multiple preset morphemes;
[0237] A fourth quantity determination module, which can be used to determine the cumulative morpheme quantity based on the historical morpheme quantity and the number of target morphemes.
[0238] In a specific embodiment, the above analysis result generation module may include:
[0239] A second result generation module, which can be used to determine that the voiceprint analysis result is the second analysis result when there is no intersection between multiple preset morphemes and the current morpheme; the second analysis result is used to indicate that it is impossible to determine whether there is a voiceprint feature of the target object in the current audio data.
[0240] In a specific embodiment, the above device may further include:
[0241] A fifth audio processing module, which can be used to perform audio processing on the current audio data based on the current processing strategy when the voiceprint analysis result is the second analysis result; the current processing strategy is the first processing strategy or the second processing strategy.
[0242] Regarding the device in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0243] Figure 5 is a block diagram of an electronic device for performing audio processing on current audio data according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as Figure 5 shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio processing method.
[0244] Figure 6 is a block diagram of another electronic device for performing audio processing on current audio data according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as Figure 6As shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements an audio processing method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0245] Those skilled in the art can understand that Figure 5 or Figure 6 the structure shown in is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0246] In an exemplary embodiment, there is also provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the audio processing method as in the embodiment of the present disclosure.
[0247] In an exemplary embodiment, there is also provided a computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, enabling the electronic device to execute the audio processing method in the embodiment of the present disclosure.
[0248] In an exemplary embodiment, there is also provided a computer program product containing instructions, when it runs on a computer, enabling the computer to execute the audio processing method in the embodiment of the present disclosure.
[0249] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0250] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0251] It can be understood that in the specific implementation manners of the present application, when it comes to data related to user information, etc., when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0252] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0253] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio processing method, characterized in that, The method includes: Obtaining current audio data corresponding to a target object and target voiceprint feature information corresponding to the target object; the target voiceprint feature information is extracted from preset recorded audio data of the target object; Performing voice detection on the current audio data to obtain a voice detection result corresponding to the current audio data; When the voice detection result indicates that the current audio data contains object audio features, performing voiceprint analysis on the current audio data based on the target voiceprint feature information to obtain a voiceprint analysis result corresponding to the current audio data; When the voiceprint analysis result indicates that none of the voiceprint feature information of the target object is included in the current audio data, performing audio processing on the current audio data based on a first processing strategy; the first processing strategy is used to indicate performing elimination processing on the current audio data.
2. The method according to claim 1, wherein The performing audio processing on the current audio data based on a first processing strategy when the voiceprint analysis result indicates that none of the voiceprint feature information of the target object is included in the current audio data includes: When the voiceprint analysis result indicates that none of the voiceprint feature information of the target object is included in the current audio data, determining an accumulated interference duration based on audio time information corresponding to the current audio data and a current interference duration reset time; When the accumulated interference duration is greater than or equal to a preset duration, performing audio processing on the current audio data based on the first processing strategy.
3. The method according to claim 2, characterized in that The determining an accumulated interference duration based on audio time information corresponding to the current audio data and a current interference duration reset time includes: When the current interference duration reset time is the starting time corresponding to the current audio data, determining the accumulated interference duration based on the audio duration of the object audio features included in the current audio data.
4. The method according to claim 2, wherein The determining an accumulated interference duration based on audio time information corresponding to the current audio data and a current interference duration reset time includes: When the current interference duration reset time is before the starting time corresponding to the current audio data, obtaining a historical interference duration of first historical audio data; the first historical audio data is audio data corresponding to the target object within a first time period, the first time period is the time period from the current interference duration reset time to the starting time corresponding to the current audio data, and the historical interference duration is the audio duration of the object audio features included in the first historical audio data; Determining the accumulated interference duration based on the historical interference duration and the audio duration of the object audio features included in the current audio data.
5. The method according to claim 2, wherein The method further includes: When the accumulated interference duration is less than the preset duration, performing audio processing on the current audio data based on a second processing strategy; the second processing strategy is used to indicate that no elimination processing is performed on the current audio data.
6. The method according to claim 1, wherein The method further includes: In the case that the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, reset the cumulative interference duration, and perform audio processing on the current audio data based on the second processing strategy.
7. The method according to claim 6, characterized in that, The method further includes: Performing noise reduction processing on the current audio data based on a preset noise template to obtain the current audio data after noise reduction; The performing voice detection on the current audio data to obtain the voice detection result corresponding to the current audio data includes: Performing voice detection on the current audio data after noise reduction to obtain the voice detection result.
8. The method according to claim 7, wherein In the case that the preset noise template does not include a preset object noise template, the step of, in the case that the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, resetting the cumulative interference duration and performing audio processing on the current audio data based on the second processing strategy includes: In the case that the voiceprint analysis result indicates that any voiceprint feature information of the target object is included in the current audio data, reset the cumulative interference duration; Performing noise reduction processing on the current audio data after noise reduction based on the target voiceprint feature information to obtain the object audio data corresponding to the target object; Performing audio processing on the object audio data corresponding to the target object based on the second processing strategy.
9. The method according to any one of claims 1-8, characterized in that The target voiceprint feature information includes the first voiceprint feature information corresponding to each of a plurality of preset morphemes; The performing voiceprint analysis on the current audio data based on the target voiceprint feature information to obtain the voiceprint analysis result corresponding to the current audio data includes: Performing speech recognition on the current audio data to obtain the current morpheme corresponding to the current audio data and the target audio data corresponding to the current morpheme; the target audio data is the valid audio data belonging to the current morpheme in the current audio data; In the case that there is an intersection between the plurality of preset morphemes and the current morpheme, performing voiceprint feature extraction processing on the target audio data to obtain the second voiceprint feature information corresponding to the current morpheme; Determining the third voiceprint feature information corresponding to the current morpheme from the first voiceprint feature information corresponding to each of the plurality of preset morphemes; Performing a matching process on the third voiceprint feature information and the second voiceprint feature information to obtain a target matching result; Generating the voiceprint analysis result based on the target matching result.
10. The method according to claim 9, wherein The target matching result includes a first matching result, and the first matching result is used to indicate that the third voiceprint feature information and the second voiceprint feature information do not match; The generating the voiceprint analysis result based on the target matching result includes: In the case that the target matching result is the first matching result, determining the cumulative morpheme quantity based on the audio time information corresponding to the current audio data and the current matching morpheme reset time. When the cumulative morpheme quantity is greater than or equal to a preset morpheme quantity, determining that the voiceprint analysis result is a first analysis result; the first analysis result is used to indicate that any voiceprint feature information of the target object is not included in the current audio data.
11. The method according to claim 10, wherein, The determining of the cumulative morpheme quantity based on the audio time information corresponding to the current audio data and the current matching morpheme reset time includes: When the current matching morpheme reset time is the starting time corresponding to the current audio data, determining, based on the current morpheme, the target morpheme quantity corresponding to the current audio data; the target morpheme quantity is the quantity of at least one morpheme included in the current morpheme among the plurality of preset morphemes; Determining the cumulative morpheme quantity based on the target morpheme quantity.
12. The method according to claim 10, wherein The determining of the cumulative morpheme quantity based on the audio time information corresponding to the current audio data and the current matching morpheme reset time includes: When the current matching morpheme reset time is before the starting time corresponding to the current audio data, obtaining the historical morpheme quantity corresponding to second historical audio data; the second historical audio data is the audio data corresponding to the target object within a second time period, the second time period is the time period from the current matching morpheme reset time to the starting time corresponding to the current audio data, the historical morpheme quantity is the quantity of at least one morpheme included in the historical morphemes corresponding to the second historical audio data among the plurality of preset morphemes, and the historical morphemes corresponding to the second historical audio data are obtained by performing speech recognition on the second historical audio data; Determining, based on the current morpheme, the target morpheme quantity corresponding to the current audio data; the target morpheme quantity is the quantity of at least one morpheme included in the current morpheme among the plurality of preset morphemes; Determining the cumulative morpheme quantity based on the historical morpheme quantity and the target morpheme quantity.
13. The method according to claim 9, wherein The performing of voiceprint analysis on the current audio data based on the target voiceprint feature information to obtain the voiceprint analysis result corresponding to the current audio data further includes: When there is no intersection between the plurality of preset morphemes and the current morpheme, determining that the voiceprint analysis result is a second analysis result; the second analysis result is used to indicate that it is impossible to determine whether the voiceprint feature of the target object exists in the current audio data.
14. The method according to claim 13, wherein The method further includes: When the voiceprint analysis result is the second analysis result, performing audio processing on the current audio data based on the current processing strategy; the current processing strategy is the first processing strategy or the second processing strategy.
15. An audio processing device, characterized in that, The apparatus includes: A current data acquisition module, configured to acquire the current audio data corresponding to the target object and the target voiceprint feature information corresponding to the target object; the target voiceprint feature information is extracted from the preset recorded audio data of the target object; A voice detection module, configured to perform voice detection on the current audio data to obtain the voice detection result corresponding to the current audio data; A voiceprint analysis module, configured to, when the voice detection result indicates that the current audio data contains object audio features, perform voiceprint analysis on the current audio data based on the target voiceprint feature information to obtain a voiceprint analysis result corresponding to the current audio data; An audio processing module, configured to, when the voiceprint analysis result indicates that none of the voiceprint feature information of the target object is included in the current audio data, perform audio processing on the current audio data based on a first processing strategy; the first processing strategy is used to indicate that elimination processing is to be performed on the current audio data.