Noise resolution method and device for assisting in on-site conversation

By collecting and processing mixed audio signals through a head-mounted device, extracting human voices using a convolutional neural network, and combining physical isolation and active noise reduction technology, the problem of clearly hearing the other person speaking in noisy environments is solved, enabling clear on-site conversations and noise reduction, and adapting to various scenarios.

CN120977329APending Publication Date: 2025-11-18NINGBO QINGSHENG TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511331110.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In noisy environments, existing technologies are not suitable for effectively hearing others speak, reducing environmental noise interference, and protecting users' ears from high-decibel noise damage.

Method used

A head-mounted device is used to collect a mixed sound signal of environmental noise and human voice. A convolutional neural network is used to identify and extract the human voice component. Combined with physical isolation and active noise reduction technology, a clear human voice signal is generated and played.

Benefits of technology

Enables clear live conversations in noisy environments, improves voice quality and communication efficiency, adapts to various noisy environments, and protects ears from noise damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977329A_ABST
    Figure CN120977329A_ABST
Patent Text Reader

Abstract

The invention discloses a noise elimination method and device for on-site conversation assistance. The method comprises the following steps: acquiring a mixed sound signal of environmental noise and human voice, performing feature extraction and analysis on the mixed sound signal, identifying noise components and human voice components in the mixed sound signal by using a neural network model, extracting the human voice components to form a human voice signal, and converting the human voice signal into human voice for playing. According to the invention, a head-mounted device integrating pickup, human voice extraction and playing functions can be realized, and conversation between people can be assisted in a noisy environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of speech or voice processing, in particular, the technical field of speech enhancement, and specifically relates to a noise elimination method and device for live conversation. BACKGROUND

[0002] With the acceleration of modern life pace and the wide application of industrial technology, people sometimes have to have a conversation in various noisy environments. Whether in a noisy street, a noisy indoor place, or a factory workshop with machine roar, environmental noise will seriously affect the quality of live conversation. Therefore, in order to protect hearing, maintain physical and mental health, and increase life experience, when we need to have a conversation in a noisy environment, we need to isolate or weaken the impact of noise on conversation while protecting the ears from high-decibel noise.

[0003] However, the existing speech or voice noise reduction technology has its limitations. The current common noise reduction methods mainly include physical isolation and active noise reduction. Physical isolation of noise mainly blocks external sound from entering through the design of earplugs or earmuffs from the physical layer. Active noise reduction technology usually uses a microphone to pick up environmental noise, and then generates a sound wave with opposite phase to the noise through the internal circuit of the earphone to cancel it out. However, both of these noise reduction technologies cannot be used in live conversation scenarios.

[0004] Firstly, using the method of physically isolating noise, the voice of the conversation object will also be isolated together with the noise, which obviously cannot be applied to live conversation scenarios.

[0005] Secondly, active noise reduction technology is applied to the scenario of online communication using earphones. In this scenario, the voice of the conversation comes from the speaker of the earphone, while the noise comes from the current environment, that is, the sources of the two are different. Using active noise reduction technology can easily pick up environmental noise and cancel it out, but this processing does not act on the voice of the conversation. Therefore, this technology cannot be applied to live conversation scenarios, because in live conversation scenarios, speech and noise are naturally combined together and received, if active noise reduction technology is used on noise containing speech, at least part of the speech will be canceled, therefore, it cannot achieve or effectively achieve the purpose of speech enhancement.

[0006] The Chinese invention patent with the authorization publication number CN114189781B proposes a noise reduction method and system for a dual-microphone neural network noise reduction earphone. The earphone collects the mixed voice of the earphone wearer (the main speaker) with noise through two microphones on the earphone, and a series of processing is performed on the voice signal to obtain a pure voice signal of the main speaker, and the voice signal is transmitted to the other party. However, this technical solution is not used in a live conversation scene, and it processes the user's own voice signal, not the voice signal of the conversation object. The existing technologies with the authorization publication numbers CN110996208B, CN114302286B, etc. all adopt similar solutions. However, these technical solutions are to make the user's conversation clear to the other party, but cannot achieve that the user can clearly hear the other party's speech or conversation.

[0007] The Chinese invention patent with the authorization publication number CN115482830B proposes a voice enhancement method and related equipment. The method obtains a noisy voice signal and target voice related data, wherein the noisy voice contains a voice signal of a target user and an interference noise signal; the target voice related data is used to indicate the voice characteristics of the target user; and the voice signal is processed by a trained voice noise reduction model according to the target voice related data to obtain a noise-reduced voice signal of the target user. However, this technical solution is to enable an intelligent device to clearly hear a person's voice in a noisy environment, i.e., to enable a machine to clearly hear and recognize a human voice. However, this solution is mainly for voice recognition, and the target user needs to register the voice signal in advance to achieve this, and cannot be applied to a live conversation scene between people. SUMMARY

[0008] (I) Technical problems to be solved

[0009] The technical problem solved by the present application is how to clearly hear the other party in a noisy environment and effectively conduct a live conversation. The present application also solves the problem of how to effectively reduce the interference of environmental noise on a live conversation in a noisy environment and protect the user's ears from high-decibel noise.

[0010] (II) Technical solutions

[0011] To solve the above technical problems, the present application proposes a noise elimination method for live conversation assistance, which is used to assist people in a conversation in a noisy environment, including the following steps: collecting a mixed sound signal of environmental noise and human voice; performing feature extraction and analysis on the mixed sound signal, identifying the noise component and human voice component in the mixed sound signal using a neural network model, extracting the human voice component to form a human voice signal; and converting the human voice signal into human voice for playing.

[0012] According to a preferred embodiment of the present application, the neural network model adopts a convolutional neural network structure, which comprises a convolutional layer for extracting local features of the sound signal and a fully connected layer for comprehensive analysis of the extracted local features, and outputs the human voice signal after noise reduction processing.

[0013] According to a preferred embodiment of the present application, the mixed sound signal is collected and the human voice is played using the same head-mounted device.

[0014] According to a preferred embodiment of the present application, the method further comprises physically isolating or actively reducing noise of the mixed sound entering the human ear using the head-mounted device.

[0015] Another aspect of the present application provides a noise elimination device for live conversation assistance, which is used to assist human-to-human conversation in a noisy environment. The device comprises: a sound pickup unit for collecting a mixed sound signal of environmental noise and human voice; a human voice extraction unit for feature extraction and analysis of the mixed sound signal, identifying noise components and human voice components in the mixed sound signal using a neural network model, and extracting the human voice components to form a human voice signal; and a playing unit for converting and playing the human voice signal as human voice.

[0016] According to a preferred embodiment of the present application, the neural network model adopts a convolutional neural network structure, which comprises a convolutional layer for extracting local features of the sound signal and a fully connected layer for comprehensive analysis of the extracted local features, and outputs the human voice signal after noise reduction processing.

[0017] According to a preferred embodiment of the present application, the noise elimination device is a head-mounted device, and when a user wears the head-mounted device, the playing unit of the head-mounted device is close to the external auditory canal of the human or close to the part of the human that can conduct sound bone.

[0018] According to a preferred embodiment of the present application, it further comprises an isolation unit arranged outside the playing unit for physically isolating the mixed sound entering the human ear.

[0019] According to a preferred embodiment of the present application, it further comprises an active noise reduction unit for generating an anti-phase signal according to the mixed sound signal; and the playing unit is further used to play the anti-phase signal of the mixed sound signal when playing the human voice.

[0020] According to a preferred embodiment of the present application, the sound pickup unit comprises one main sound pickup unit and two auxiliary sound pickup units; the main sound pickup unit is located close to the mouth of the user; and the two auxiliary sound pickup units are respectively located on both sides of the head-mounted device.

[0021] According to a preferred embodiment of the present application, the noise elimination device further comprises a talk processing unit, which is used to convert the sound picked up by the main pickup unit into a sound signal for transmission, and transmit the sound signal received from outside to the playing unit for playing.

[0022] According to a preferred embodiment of the present application, the noise elimination device further comprises a master control unit, which controls the overall device to switch between the live conversation mode and the talk mode by controlling the operation of the talk processing unit and the voice extraction unit.

[0023] According to a preferred embodiment of the present application, the noise elimination device further comprises a preprocessing unit, wherein the preprocessing unit performs directional enhancement processing on the sound signals received by the main pickup unit and the two auxiliary pickup units to strengthen the voice signal of the main pickup unit, and then transmits the voice signal of the main pickup unit to the talk processing unit; the preprocessing unit also performs reverse enhancement processing on the sound signals received by the main pickup unit and the two auxiliary pickup units to weaken or eliminate the sound signal of the main pickup unit before transmitting it to the voice extraction unit, so as to extract the voice other than the user's own voice.

[0024] (III) Beneficial Effects

[0025] Compared with the prior art, the present application has the following beneficial effects:

[0026] 1. Improved noise reduction effect: Compared with traditional earphones, the method and device of the present application can ensure communication efficiency and achieve stronger speech noise reduction algorithm, thereby better eliminating various non-stationary noises in complex noisy environments, creating a quiet auditory environment for users, and protecting the ears from noise damage.

[0027] 2. Realize clear live conversation: The present application uses neural network technology to accurately extract and enhance voice, which significantly improves the intelligibility of the speech heard by the other party during real-time communication in a noisy environment, and the intelligibility is close to the level of speaking in a quiet environment, effectively avoiding communication interruption, misunderstanding and other problems caused by noise interference, and greatly improving the speech quality and communication efficiency.

[0028] 3. Adapt to multiple scenarios: Since the neural network model is trained on a large amount of different scene noise data, the earphone of the present application can quickly adapt to various noisy environments, and can stably play the noise reduction and real-time communication enhancement function in various scenarios such as outdoor traffic noise, indoor noisy voice, and industrial environment noise, and has wide applicability. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 A schematic diagram of an application scenario of the present application. The diagram shows the scenario of factory internal training.

[0030] Figure 2 This is a schematic diagram illustrating another application scenario of the present invention. The diagram shows a scene of conversation taking place inside a factory building.

[0031] Figure 3 This is a schematic diagram of one embodiment of the noise reduction device for assisting on-site conversation according to the present invention.

[0032] Figure 4 This is an internal functional unit architecture diagram of an embodiment of the noise cancellation device for assisting on-site conversation according to the present invention.

[0033] Figure 5 This is an internal functional unit architecture diagram of an embodiment of the noise cancellation device for assisting on-site conversation according to the present invention. Figure 6 This is a schematic diagram of another embodiment of the noise reduction device for assisting on-site conversation according to the present invention. Figure 7 yes Figure 6 A schematic diagram of an embodiment of the internal functional unit architecture of a noise cancellation device for assisting in on-site conversation. Figure 8 yes Figure 6 A schematic diagram of another embodiment of the internal functional unit architecture of a noise cancellation device for assisting in on-site conversation. Detailed Implementation

[0034] To address the aforementioned technical problems, this invention proposes a noise reduction method specifically for assisting in on-site conversations, designed to facilitate normal conversation between people in noisy environments such as factories, streets, and conference rooms. On-site conversation in this invention refers to face-to-face speech between humans, excluding dialogue between humans and machines, and also excluding wireless or long-distance voice transmission via electrical signals, infrared light, electromagnetic waves, etc. In other words, phone calls, mobile phone calls, or internet communication are not considered on-site conversations as defined in this invention. However, the on-site conversation assistance function of this invention can also be integrated with other functions into a single device; for example, the on-site conversation assistance function of this invention can be added to existing headsets used for making calls.

[0035] Figure 1This is a schematic diagram illustrating an application scenario of the present invention. The diagram shows a training session within a factory. Conducting technical training within a factory often presents challenges because the operation of machinery frequently generates loud noise, which can mask the trainer's voice and significantly impact the quality of the training. A traditional solution is to use a loudspeaker, but excessive volume can impair hearing. Another solution is to utilize wireless signal transmission technology, which requires both the trainer and trainees to be equipped with headphones that support wireless signal transmission. These headphones transmit the trainer's voice to the trainees. However, this method requires establishing a one-to-many wireless connection between the headphones or using networking equipment for a small-scale network. For large numbers of people, the operation is overly complex and prone to connection failures. Furthermore, wireless transmission involves signal modulation, demodulation, and encoding / decoding, which often introduces significant latency, affecting the user experience.

[0036] The present invention adopts a completely different approach. The present invention directly receives a mixed sound signal containing noise and human voice on site, and directly extracts and plays the human voice from the mixed sound signal.

[0037] like Figure 1 As shown, each student can be provided with an earpiece of this invention. This earpiece does not require communication with any other device and is therefore ready to use immediately. The earpiece uses its onboard microphone to directly receive a mixed sound signal of noise and human voice. The processor inside the earpiece processes the signal, extracts the human voice signal, and plays it through a speaker. When the instructor is lecturing, students can clearly hear the instructor's voice, and even the voices of other students, through the earpiece. If the instructor does not need to interact with students during the lecture (i.e., does not need to hear the students' voices), they can choose not to wear the earpiece. Conversely, when interaction with students is required, the instructor can wear the earpiece to hear their voices.

[0038] In order to extract and analyze features from the mixed audio signal, the present invention preferably uses a neural network model to identify the noise component and the human voice component in the mixed audio signal, and extracts the human voice component to form a human voice signal.

[0039] It should be noted that headphones are a typical implementation of this invention, but this invention does not exclude any other devices with sound pickup and playback functions. For example, some wearable devices also have sound pickup and playback functions, such as glasses, hats, and helmets. However, this invention prefers to implement it as a head-mounted device. The head-mounted device in this invention includes any device that is fixed to a person's head in any way, including ear-hook, in-ear, neckband, headband, etc. Furthermore, the headphones can be not only traditional air conduction headphones but also bone conduction headphones.

[0040] The pickup unit adopted by the present application is preferably a microphone with high sensitivity and wide frequency response characteristics, so as to be able to accurately capture sound signals of various frequencies.

[0041] In addition, the earphone of the present application can also be combined with traditional physical isolation or active noise reduction technology. On the one hand, the earphone can be equipped with physical isolation elements, and high-quality sound insulation materials such as high-density silica gel earplugs or ergonomic ear covers are used in the structural design to reduce the transmission of external noise into the ear through physical means as much as possible. On the other hand, traditional active noise reduction technology can also be combined, and according to the mixed sound signals received by the microphone, a counter signal is generated, which is output through the earphone speaker and cancels out the residual noise in the ear canal, further reducing the noise level.

[0042] Figure 2 The figure shows the scene of talking in a factory building. Different from the previous scene, in this scene, the personnel talking and the place of talking may not be fixed. For example, the scene of the leader visiting or receiving visitors, the personnel talking may be random, the way of talking may be very flexible, it may be one-to-one, one-to-many, many-to-many, and may be switched at any time. In this case, if the traditional wireless communication method is used to realize the transmission of sound, the process of establishing a communication connection will be very complex, and it is almost impossible to realize in a quick and effective way. Using the present application, you only need to wear an earphone to communicate easily, without considering the changes of the personnel and the place of talking.

[0043] In order to make the purpose, technical scheme and advantages of the present application clearer and more apparent, the present application will be further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0044] Figure 3 is a structural schematic diagram of an embodiment of the noise elimination device for on-site conversation assistance of the present application. As shown in Figure 3 , the noise elimination device is realized as an earphone, which has two microphones 1 (only one side is shown in the figure) for picking up sound. In this embodiment, a MEMS microphone with a frequency response range of 20Hz-20kHz is selected as the pickup element. The two microphones 1 are respectively arranged on one side of the outer side of the earphone, so as to collect mixed sound signals of environmental noise and human voice in all directions.

[0045] In other embodiments, in addition to using MEMS microphones, electret microphones can also be considered. Electret microphones have the characteristics of low cost and high sensitivity, and the frequency response range can also meet the needs of the present application. In some cost-sensitive product applications, electret microphones can be used as an alternative to MEMS microphones. In this embodiment, the number of microphones is two and is arranged on the side, but the number of microphones can also be one, three or more, and the position can also be changed according to the actual design needs, as long as it can receive the surrounding sound.

[0046] In this embodiment, the earphone includes earplugs and ear covers. The earplugs are made of soft and sound-insulating silicone material, and the sound insulation volume is between 20dB and 30dB. The ear cover part is filled with high-density sponge and covered with sound-insulating fabric on the surface to further enhance the sound insulation effect.

[0047] In other embodiments, for the earplug part, in addition to silicone material, memory sponge material can also be selected. Memory sponge has good fit, can automatically adjust according to the shape of the ear canal, and provides good sound insulation effect, and the sound insulation volume can reach about 25dB. In some application scenarios with high requirements for wearing comfort, memory sponge earplugs may be more popular. For the ear cover part, carbon fiber composite material can be used instead of the traditional sponge and fabric combination. Carbon fiber composite material has the characteristics of light weight, high strength and good sound insulation performance, which can reduce the overall weight of the earphone while ensuring sound insulation effect, and improve the comfort and portability of the user.

[0048] Figure 4 is Figure 3 a schematic diagram of one embodiment of the internal functional unit architecture of the noise cancellation device for live conversation assistance. As Figure 4 shown, the device of this embodiment includes a main pickup unit, two auxiliary pickup units, a human voice extraction unit, a left playback unit and a right playback unit. The pickup unit can be implemented as a microphone or other pickup element. The human voice extraction unit can be implemented by any element with corresponding data processing capability, such as a microcontroller (MCU), a dedicated audio DSP, a neural network processor (NPU), etc. The left and right playback units can use traditional loudspeakers, or moving coil units or piezoelectric ceramics for bone conduction.

[0049] It should be noted that in other embodiments, the number and position of the pickup units and playback units in the above embodiments can be changed according to the overall structure and purpose of the device.

[0050] In this embodiment, the human voice extraction unit is implemented using a neural network processor (NPU), and a deep neural network model is built using the Pytorch framework. The input layer of the model receives the mixed sound signal after amplification and preprocessing, with a signal sampling rate of 16 kHz and a quantization accuracy of 16 bits. The convolutional layer is set with 3 convolutional kernels, with kernel sizes of 3x3, 5x5, and 7x7, and a step size of 1. The features of the sound signal are extracted through multiple convolution operations. The pooling layer uses the maximum pooling method, with a pooling window size of 2x2 and a step size of 2, to reduce the dimensionality of the feature map after convolution. The fully connected layer contains 256 neurons and uses the ReLU activation function, and outputs the human voice signal and noise estimation signal after noise reduction processing. A large amount of audio data containing various environmental noise (such as street traffic noise, subway noise, restaurant noise, etc.) and different human voices (including human voices of different genders, ages, and accents) is used to train the model, with a training duration of 1000 hours. The Adam optimizer is used as the optimization algorithm, and the learning rate is set to 0.001.

[0051] In other embodiments, in addition to using a convolutional neural network (CNN) structure, a recurrent neural network (RNN) or its variants such as a long short-term memory network (LSTM) can also be used. RNN and LSTM models have an advantage in processing data with time series characteristics, and may be able to better capture the dynamic characteristics of continuously changing sound signals, thereby achieving different effects in noise reduction and human voice enhancement. For example, in some scenarios where real-time performance is required, the LSTM model may be able to process sound signals faster and adjust the noise reduction and enhancement strategy in a timely manner.

[0052] The above embodiments of the present application use a neural network module to enhance and extract human voices from the outside world, remove noise, retain clear human voices, and enhance the fullness of human voices. In cooperation with other modules of the earphone, a quiet simulated environment is constructed, and clear communication is carried out, thereby improving the quality of communication and communication efficiency.

[0053] In addition, in addition to physically isolating sound, the present application can also be combined with active noise reduction technology. The active noise reduction circuit generates an anti-phase noise signal through a digital signal processor (DSP) based on the noise signal picked up by the microphone, and outputs it through a power amplifier after amplification by the earphone speaker. The response time of the active noise reduction circuit is less than 5ms, and it can quickly respond to changes in noise.

[0054] Figure 5 is another embodiment of the internal functional unit architecture of the noise elimination device for live conversation assistance of the present application. Compared with Figure 4In different embodiments, the device in this embodiment further includes an active noise reduction unit, which generates an inverted signal based on the mixed sound signal obtained by the main pickup unit and the auxiliary pickup unit, and sends it to the playback unit. When playing human voice, the playback unit is also used to play the inverted signal of the mixed sound signal.

[0055] Compared to traditional headphones, the headphones of this invention can implement a more powerful noise reduction algorithm because they can ensure communication efficiency. In complex and noisy environments, they can better eliminate various non-stationary noises, creating a quiet listening environment for users and protecting their ears from noise damage.

[0056] Figure 6 This is a schematic diagram of another embodiment of the noise cancellation device for assisting on-site conversation according to the present invention. Compared to Figure 3 In the previous embodiment shown, the headphones also include a microphone 2 for picking up the user's own voice. For ease of distinction, microphone 2 can be referred to as the main microphone, while the two microphones 1 located on the side can be referred to as auxiliary microphones (only one side is shown in the figure). The main microphone 2 extends forward via a linkage, close to the user's mouth, to facilitate clear pickup of the user's own voice.

[0057] The headset in this embodiment can be seen as adding the on-site conversation assistance function of the present invention to a conventional headset used for calls. That is, the main microphone 2 added in the embodiment can pick up the user's own voice during a call, while the auxiliary microphone is only used to pick up ambient sounds during on-site conversations.

[0058] It should be noted that in this invention, "call" refers to real-time remote communication between people, but is not limited to the medium of the call; it can be a telephone or a call transmitted over the Internet.

[0059] Figure 7 yes Figure 6 A schematic diagram of an embodiment of the internal functional unit architecture of a noise cancellation device for assisting in on-site conversation.

[0060] like Figure 7 As shown, compared to Figure 4 The architecture of the device also includes a call processing unit and a main control unit. The function of the call processing unit is to convert the sound picked up by the main pickup unit into an audio signal for transmission, and simultaneously send externally received audio signals to the playback unit for playback. Therefore, the call processing unit should have signal transmission and reception capabilities as well as signal encoding and decoding functions. In this embodiment, the main pickup unit is only connected to the call processing unit and not to the voice extraction unit.

[0061] In addition, in order to switch the earphone between the live conversation mode and the talk mode, the master control unit can be used to coordinate the operation of the units, especially the operation of the voice extraction unit and the talk processing unit. The master control unit can be realized by a component with digital signal processing capability. The master control unit can switch between the live conversation mode and the talk mode automatically or according to the user's operation.

[0062] For automatic switching, when the talk processing unit receives a talk request, it reports the talk request state to the master control unit, at which time the master control unit stops the operation of the voice extraction unit and enables the operation of the talk processing unit. When the talk processing unit ends the talk, it reports the end of the talk state to the master control unit, at which time the master control unit stops the operation of the talk processing unit and enables the operation of the voice extraction unit.

[0063] For manual switching, an operation component such as a talk button, a key, a pull rod, a pull disc, etc. can be arranged at a proper position of the earphone (for example, on the microphone support rod). Thus, the user can manually operate the operation component to switch between the live conversation mode and the talk mode, and the master control unit can determine which working mode to switch to by detecting the current state of the operation component.

[0064] As an alternative to the above embodiment, the master control unit, the voice extraction unit and the talk processing unit can be combined into one unit, or part of the functions of the above units can be combined into one unit. In some specific embodiments, the voice extraction unit and the talk processing unit do not need to be switched, and thus the master control unit can be omitted.

[0065] Figure 8 is Figure 6 Another embodiment of the internal functional unit architecture of the live conversation auxiliary noise elimination device. The biggest difference between this embodiment and the embodiment shown in Figure 7 The biggest difference between this embodiment and the embodiment shown in

[0066] Figure 7 In the embodiment of the application, the voice signal received by the talk processing unit from the main pickup unit can contain a large amount of noise signal when in a noisy environment, which will greatly reduce the quality of the talk. Therefore, in this embodiment, the pre-processing unit performs directional enhancement processing on the sound signals received by the three microphones to strengthen the voice signal of the main pickup unit, and then sends the voice signal of the main pickup unit to the talk processing unit.

[0067] Meanwhile, the preprocessing unit also performs reverse enhancement processing on the sound signals received by the three microphones to weaken or eliminate the sound signals of the main pickup unit, and then sends the signals to the voice extraction unit, so as to extract the voice of the user except the user's own voice, and thus can avoid enhancing the user's own voice and improve the use experience when the user and others talk at the same time in a live conversation scene.

[0068] The directional enhancement processing and reverse enhancement processing techniques can adopt existing sound signal processing techniques, which will not be described here.

[0069] The present application can significantly improve the intelligibility of the voice heard by the other party in a noisy environment, and the intelligibility is close to the level of conversation in a quiet environment, effectively avoiding the problems of conversation interruption and misunderstanding caused by noise interference, and greatly improving the conversation quality and communication efficiency.

[0070] Since the neural network model is trained by a large amount of different scene noise data, the earphone of the present application can quickly adapt to various noisy environments, and can stably play the noise reduction and real-time conversation enhancement function in various scenes such as outdoor traffic noise, indoor noisy voice, and industrial environment noise, and has wide applicability.

[0071] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application, and it should be understood that the above description is only for specific embodiments of the present application and is not intended to limit the present application, and any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A noise cancellation method for on-the-spot conversation assistance for assisting a conversation between people in a noisy environment, characterized by, The method comprises the following steps: collecting a mixed sound signal of environmental noise and human voice; extracting and analyzing features of the mixed sound signal, identifying noise components and human voice components in the mixed sound signal by using a neural network model, and extracting the human voice components to form a human voice signal; converting the human voice signal into human voice for playing.

2. The noise-cancelling method for live conversation assistance of claim 1, wherein, The neural network model adopts a convolutional neural network structure, which comprises a convolutional layer and a fully connected layer, the convolutional layer is used to extract local features of a sound signal, and the fully connected layer is used to comprehensively analyze the extracted local features and output a human voice signal after noise reduction processing.

3. The noise-cancellation method for live conversation assistance of claim 1 or 2, wherein, The mixed sound signal is collected and the human voice is played by using the same head-mounted device.

4. The noise-cancellation method for live conversation assistance of claim 3, wherein, Further comprising: physically isolating or actively reducing noise of the mixed sound entering the human ear by using the head-mounted device.

5. A noise-cancellation apparatus for live conversation assistance for assisting a conversation between people in a noisy environment, characterized in that, comprising: a sound pickup unit for collecting a mixed sound signal of environmental noise and human voice; a human voice extraction unit for extracting and analyzing features of the mixed sound signal, identifying noise components and human voice components in the mixed sound signal by using a neural network model, and extracting the human voice components to form a human voice signal; a playing unit for converting the human voice signal into human voice for playing.

6. The noise-cancelling device for live conversation assistance of claim 5, wherein, The neural network model adopts a convolutional neural network structure, which comprises a convolutional layer and a fully connected layer, the convolutional layer is used to extract local features of a sound signal, and the fully connected layer is used to comprehensively analyze the extracted local features and output a human voice signal after noise reduction processing.

7. Noise cancellation apparatus for on-site conversation assistance according to claim 5 or 6, characterized in that, The noise elimination device is a head-mounted device, when a user wears the head-mounted device, the playing unit of the head-mounted device is close to the external auditory canal of the human or close to the part of the human that can conduct sound bone conduction.

8. The noise-cancelling device for live conversation assistance of claim 7, wherein, Further comprising an isolation unit arranged outside the playing unit, for physically isolating the mixed sound entering the human ear.

9. The noise-cancelling device for live conversation assistance of claim 7, wherein, Further comprising an active noise reduction unit for generating an anti-phase signal according to the mixed sound signal; The playing unit also plays the anti-phase signal of the mixed sound signal when playing the human voice.

10. The noise elimination device for live conversation assistance according to claim 7, wherein the sound pickup unit comprises a main sound pickup unit and two auxiliary sound pickup units; the main sound pickup unit is located close to the mouth of the user; the two auxiliary sound pickup units are respectively located on both sides of the head-mounted device.

11. The noise-cancelling device for live conversation assistance of claim 10, wherein, The noise elimination device further comprises a call processing unit for converting the sound picked up by the main sound pickup unit into a sound signal for sending, and sending the sound signal received from the outside to the playing unit for playing.

12. The noise-cancelling device for live conversation assistance of claim 11, wherein, The noise elimination device further comprises a master control unit for controlling the overall device to switch between the live conversation mode and the call mode by controlling the working of the call processing unit and the human voice extraction unit.

13. The noise-cancelling device for live conversation assistance of claim 11, wherein, The noise elimination device further comprises a preprocessing unit, wherein the preprocessing unit performs directional enhancement processing on the sound signals received by the main sound pickup unit and the two auxiliary sound pickup units to strengthen the human voice signal of the main sound pickup unit, and then sends the human voice signal of the main sound pickup unit to the call processing unit. The preprocessing unit also performs reverse enhancement processing on the sound signals received by the main pickup unit and the two auxiliary pickup units to attenuate or eliminate the sound signals of the main pickup unit before sending them to the voice extraction unit, so as to extract voices other than the user's own voice.

Citation Information

Patent Citations

  • A wireless earphone and its noise reduction method

    CN110996208B

  • Noise reduction method and system for dual-microphone neural network noise reduction headphones

    CN114189781B

  • A method, device, equipment and storage medium for reducing call noise

    CN114302286B

  • Speech enhancement methods and related equipment

    CN115482830B