Audio processing method, device and system, storage medium and computer program product
By acquiring the speaker's location information in real time and dynamically adjusting the sound screen range, the problem of decreased sound pickup quality caused by the speaker moving away from the sound screen was solved, achieving high-quality audio processing effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, the speaker may move out of the soundstage, resulting in a decrease in sound pickup quality, an inability to effectively filter noise, and a failure to focus the speaker's voice.
By acquiring the speaker's location information in real time, the audio screen range is dynamically adjusted to ensure that the speaker is always within the audio screen. Noise signals are eliminated using microphones and processors, and the audio screen range is determined using image signal and sound source localization technology.
It improves the sound pickup quality, ensuring the speaker's voice is clear and natural, effectively filters noise, and enhances the audio effect.
Smart Images

Figure CN121865157A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to an audio processing method, apparatus, system, storage medium, and computer program product. Background Technology
[0002] In audio and video conferencing scenarios, audio quality directly impacts meeting effectiveness. High-quality sound pickup ensures the speaker's voice is clear and natural, better supporting remote collaboration and improving meeting efficiency. This is especially true in complex meeting environments, such as open-plan offices or large conference rooms, where there is often ambient noise, echoes, reverberation, and other noise such as voices not belonging to the speaker. Good sound pickup technology effectively filters noise, focuses the speaker's voice, and provides clear audio.
[0003] Currently, intelligent sound barrier technology can be used to improve sound pickup quality. Intelligent sound barrier technology works by creating a virtual acoustic barrier (i.e., a sound screen) to block sound interference from outside the sound screen's range. For example, audio signals can be collected using microphones in conference terminal equipment, and the processor in the conference terminal equipment can eliminate sounds outside the sound screen's range from the audio signal, thereby improving the sound quality of the audio signal within the sound screen's range.
[0004] In related technologies, the audio censorship range is preset, i.e., pre-configured. For example, based on experience, the speaker is usually positioned within a 120-degree angle and 6 meters directly in front of the conference terminal equipment. Therefore, this 120-degree angle and 6-meter area can be pre-configured as the audio censorship range. However, since the speaker may still enter outside the audio censorship range, in this case, the related technology will also cancel out the speaker's voice, resulting in a significant decrease in sound pickup quality. Summary of the Invention
[0005] This application provides an audio processing method, apparatus, system, storage medium, and computer program product that ensures the speaker remains within the soundstage, thereby improving sound pickup quality. The technical solution is as follows:
[0006] Firstly, an audio processing method is provided, the method comprising:
[0007] Acquire the initial audio signal, which includes the speaker's audio signal and noise signal; determine the speaker's location information; determine the audio censorship range based on the speaker's location information; and eliminate the noise signal in the initial audio signal based on the audio censorship range.
[0008] In other words, based on the speaker's real-time location information, the audio screen range is dynamically determined, so that the audio screen range can change with the speaker's position, thereby ensuring that the speaker is always within the audio screen range and improving the sound pickup quality.
[0009] In one possible implementation, the speaker's location information includes location information at a first moment and location information at a second moment, with the first moment preceding the second moment. Determining the audio-visual censorship range based on the speaker's location information includes: if the speaker's location changes based on the first and second moment location information, then determining the audio-visual censorship range based on the second moment location information. In other words, the audio-visual censorship range is updated when the speaker's location changes.
[0010] In another possible implementation, the speaker's location information includes the location information of the current frame; based on the speaker's location information, the audio-visual scope is determined, including: determining the audio-visual scope based on the location information of the current frame. That is, the audio-visual scope is determined in real time.
[0011] In one possible implementation, the speaker's location information includes the speaker's azimuth angle and / or position coordinates, which indicate the speaker's direction relative to the target microphone. The audio veil range includes the area pointed to by this direction, and the target microphone is the microphone that collects the speaker's audio signal. The audio veil in this implementation can be called a fan-shaped audio veil.
[0012] The sound screen range includes the area between the first ray and the second ray, which intersect at the target microphone, with the direction between the first ray and the second ray.
[0013] In another possible implementation, the speaker's location information includes the speaker's location coordinates, and the audio-visual screen range includes the area centered on those coordinates. This type of audio-visual screen can be called a fixed-point audio-visual screen.
[0014] In one possible implementation, the sound screen encompasses a region centered at the given location coordinates. This type of sound screen can be called a circular fixed-point sound screen.
[0015] In one possible implementation, determining the speaker's location information includes: acquiring an image signal that includes image information of the speaker; and determining the speaker's location information based on the image signal. That is, target detection and tracking of the speaker can be performed based on the image signal, and the audio-visual range can be dynamically updated based on the speaker's real-time location information.
[0016] In one possible implementation, the initial audio signal includes audio signals collected by multiple microphones, where the spectral range is the spectral range corresponding to the target microphone, and the target microphone is one of the multiple microphones. Based on the spectral range, noise signals in the initial audio signal are eliminated, including: based on the spectral range, noise signals in the audio signal collected by the target microphone are eliminated. That is, in a scenario with multiple microphones, the audio signal collected by one of the microphones (referred to as the target microphone) is processed and played back.
[0017] In one possible implementation, the target microphone is the microphone closest to the speaker among the multiple microphones. That is, the closer the microphone is to the speaker, the higher the quality of the audio signal it captures. Playing the audio signal captured by the microphone closest to the speaker can maximize the quality of the final played audio signal.
[0018] In one possible implementation, the multiple microphones are distributed across multiple terminal devices, with each terminal device including at least one microphone. The method further includes: synchronizing the speaker's location information to the multiple terminal devices; receiving distance information sent by the multiple terminal devices respectively, the distance information representing the distance between the respective terminal device and the speaker; and determining a target microphone from the multiple microphones based on the distance information sent by the multiple terminal devices.
[0019] In another possible implementation, the multiple microphones are distributed across multiple terminal devices, with each terminal device including at least one microphone. The method further includes: determining the distances between the speaker and each of the multiple terminal devices based on the location information of the multiple terminal devices and the location information of the speaker; and determining the target microphone from the multiple microphones based on the distances between the speaker and each of the multiple terminal devices.
[0020] In one possible implementation, the number of speakers is M, where M is a positive integer; and when M is greater than 1, the audio / video range includes the audio / video ranges corresponding to each of the M speakers. That is, the above method supports scenarios with multiple speakers.
[0021] Secondly, an audio processing apparatus is provided, which has the function of implementing the audio processing method described in the first aspect. The audio processing apparatus includes one or more modules for implementing the audio processing method provided in the first aspect.
[0022] Thirdly, another audio processing apparatus is provided, the audio processing apparatus including a processor and a memory; the memory is used for a computer program, and the processor is used to execute the computer program to implement the audio processing method provided in the first aspect above.
[0023] Fourthly, a conference system is provided, the conference system including a microphone and a processor, the microphone being used to acquire an initial audio signal, and the processor being used to execute the steps of the audio processing method provided in the first aspect above.
[0024] Fifthly, a computer device is provided, comprising a processor and a memory, the memory being used to store a program for executing the audio processing method provided in the first aspect, and to store data related to implementing the audio processing method provided in the first aspect. The processor is configured to execute the program stored in the memory.
[0025] In one possible implementation, the computer device may further include a communication bus for establishing a connection between the processor and the memory.
[0026] In a sixth aspect, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the steps of the audio processing method described in the first aspect.
[0027] In a seventh aspect, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the steps of the audio processing method described in the first aspect.
[0028] The technical effects achieved by the second to seventh aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;
[0030] Figure 2 This is a flowchart of an audio processing method provided in an embodiment of this application;
[0031] Figure 3 This is a schematic diagram of a fan-shaped sound screen provided in an embodiment of this application;
[0032] Figure 4 This is a schematic diagram of a circular fixed-point sound screen provided in an embodiment of this application;
[0033] Figure 5 This is a schematic diagram of a multi-person audio-visual subtitle provided in an embodiment of this application;
[0034] Figure 6 This is a schematic diagram of another multi-person audio-visual subtitle provided in an embodiment of this application;
[0035] Figure 7This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application;
[0036] Figure 8 This is a schematic diagram of another audio processing device provided in the embodiments of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0038] First, some terms used in the embodiments of this application will be introduced.
[0039] Object detection: In image scenes, object detection typically refers to detecting targets of interest (also called objects), such as human bodies or faces, determining their location, or even identifying their category. In the embodiments of this application, object detection may include face detection and human body detection with specified behaviors.
[0040] Object tracking: In image scenes, object tracking typically refers to detecting one or more objects of interest in an image and tracking their trajectories. When the number of objects is unknown, multiple objects of interest can be detected in the image, each assigned an identity document (ID), and then their trajectories can be tracked.
[0041] Sound source location (SSL): This method uses multiple microphones at different locations in the environment to collect sound signals. Since the sound signals arrive at each microphone at different times with varying degrees of delay, algorithms (such as the time difference of arrival (TDOA) algorithm) are used to process the sound signals collected by the multiple microphones. This process yields the direction and distance of the sound source relative to one or more microphones. The direction can be represented by azimuth angle, pitch angle, etc., and the distance can be represented by position coordinates.
[0042] Intelligent sound barrier: It creates a virtual acoustic barrier, or sound screen, by using sound source localization technology to eliminate or filter sound signals outside the range of the sound screen.
[0043] The implementation environment of the embodiments of this application will be described next.
[0044] The audio processing method provided in this application can be applied to various audio and video scenarios, such as conference scenarios, teaching scenarios, and training scenarios. These scenarios can be divided into single-person scenarios and multi-person scenarios based on the number of speakers. A single-person scenario refers to a scenario where there is only one speaker within the same time period. In this scenario, a corresponding audio-visual subtitle range can be determined for this single speaker. A multi-person scenario refers to a scenario where there can be more than one speaker within the same time period. When the number of speakers is greater than one, corresponding audio-visual subtitle ranges can be determined for each of the multiple speakers.
[0045] In this embodiment, the audio-visual scene (also known as an audio-visual system, such as a conferencing system) includes a microphone (or microphone array) and a processor. The microphone and processor establish a communication connection. The microphone is used to acquire initial audio signals, and the processor is used to dynamically determine the soundstage range according to the audio processing method provided in this embodiment, and process the initial audio signal based on the soundstage range.
[0046] The number of microphones can be one or more. We will discuss these scenarios in detail below.
[0047] When there is only one microphone, it can be located in a terminal device, meaning the audio / video scene includes a terminal device. Based on this, the processor can be a processor within the terminal device, or a processor within a control device (such as a central control device), or a processor in a server; that is, the audio / video scene may also include a control device and / or a server. This application does not limit the specific form, structure, or type of the processor.
[0048] In some embodiments, the audio-visual scene also includes a camera, which establishes a communication connection with the processor. The camera is used to acquire image signals, and the processor can also be used to detect the speaker based on the image signals. Alternatively, in one possible implementation, the camera can also detect the speaker based on the image signals and send the detection result to the processor; that is, the camera integrates a processor that can be used for speaker detection. The camera can be a camera within the aforementioned terminal device, or it can be located outside the terminal device; that is, the camera can be a relatively independent device. The camera can be placed at any suitable location in the audio-visual scene. This application does not limit the specific shape, structure, or type of the camera.
[0049] When there are multiple microphones, these microphones can be distributed across one or more terminal devices. Each terminal device includes at least one microphone. In other words, the audio-visual scene may include one or more terminal devices.
[0050] In the case where multiple microphones are distributed across a single terminal device, the processor can be a processor within the terminal device, or a processor within a control device (such as a central control device), or a processor within a server. In other words, the audio-visual scenario may also include a control device and / or a server. This application does not limit the specific form, structure, or type of the processor. Furthermore, in some embodiments, the audio-visual scenario also includes a camera. This camera may have the same function, distribution, specific form, structure, and type as the camera in the case of a single microphone, and will not be described further here.
[0051] In a scenario where multiple microphones are distributed across multiple terminal devices, these terminal devices can be located at different positions to collect the speaker's audio signal from different locations. The processor can be a processor in a first terminal device among the multiple terminal devices. The first terminal device can be any one of the multiple terminal devices; that is, the first terminal device can collect the audio signals collected by the multiple microphones and process them according to the audio processing method provided in the embodiments of this application. Alternatively, the processor can be a processor in a control device or a processor in a server; that is, the audio-visual scenario may also include a control device and / or a server. The embodiments of this application do not limit the specific form, structure, type, etc., of the processor.
[0052] In some embodiments, the audio-visual scene further includes one or more cameras. When the audio-visual scene includes one camera, the camera can be a camera on a second terminal device among the plurality of terminal devices. The second terminal device can be any one of the plurality of terminal devices, and the second terminal device and the first terminal device can be the same terminal device or different terminal devices. The camera can also be a relatively independent imaging device. The imaging device can be arranged at any possible location in the audio-visual scene. When the audio-visual scene includes multiple cameras, any one of the multiple cameras can be a camera on any terminal device, or it can be a relatively independent imaging device; that is, some or all of the terminal devices can also include at least one camera. The multiple cameras can be arranged at any possible location in the audio-visual scene. This application embodiment does not limit the specific shape, structure, type, etc., of the camera.
[0053] The server in the above embodiments can be a single server or a cloud service cluster, and this application does not limit this.
[0054] In the embodiments of this application, the microphone can be a one-dimensional microphone array or a multi-dimensional microphone array, such as a cross-shaped microphone array. The embodiments of this application do not limit the structure, type, or pickup range of the microphone.
[0055] The aforementioned processor can be a general-purpose central processing unit (CPU), a natural network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solutions of this application, such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. In some embodiments, the aforementioned PLD is a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0056] The aforementioned terminal device can be a sound pickup device or a device with more integrated functions; this application does not limit this.
[0057] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device can be any of the terminal devices mentioned above, or the control devices or servers mentioned above. The computer device includes one or more processors 101, a communication bus 102, a memory 103, and one or more communication interfaces 104.
[0058] The processor 101 can be any of the processors mentioned above, which will not be repeated here.
[0059] The communication bus 102 is used to transmit information between the aforementioned components. In some embodiments, the communication bus 102 is divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 1 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0060] In some embodiments, the memory 103 may be a read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), optical disc (including compact disc read-only memory (CD-ROM), compressed optical disc, laser disc, digital versatile optical disc, Blu-ray disc, etc.), magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 103 exists independently and is connected to the processor 101 via the communication bus 102, or the memory 103 is integrated with the processor 101.
[0061] Communication interface 104 uses any transceiver-like device for communicating with other devices or communication networks. Communication interface 104 includes a wired communication interface, and in some embodiments, also includes a wireless communication interface. The wired communication interface is, for example, an Ethernet interface. In some embodiments, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.
[0062] In some embodiments, multiple processors are included, such as Figure 1 The processors 101 and 105 shown are illustrated. Each of these processors is either a single-core processor or a multi-core processor. In one possible implementation, a processor here refers to one or more devices, circuits, and / or processing cores used for processing data (such as computer program instructions).
[0063] In a specific implementation, as one embodiment, an output device 106 and an input device 107 are also included. The output device 106 communicates with the processor 101 and can display information in various ways. For example, the output device 106 can be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device 107 communicates with the processor 101 and can receive user input in various ways. For example, the input device 107 can be a microphone, a camera, a mouse, a keyboard, a touchscreen device, or a sensing device.
[0064] In some embodiments, memory 103 stores program code 110 for executing the scheme of this application, and processor 101 is capable of executing the program code 110 stored in memory 103. The program code includes one or more software modules, which, through processor 101 and the program code 110 in memory 103, can implement the following... Figure 2 The audio processing method provided in the embodiment.
[0065] As described above, embodiments of this application also provide a conference system. In some embodiments, the conference system includes a terminal device and a control device, with a wireless or wired communication connection established between the terminal device and the control device. The terminal device is used to acquire an initial audio signal, and the control device is used to process the initial audio signal according to the audio processing method provided in the embodiments of this application. In other embodiments, the conference system includes a terminal device and a server. The terminal device is used to acquire the initial audio signal, and the server is used to process the initial audio signal according to the audio processing method provided in the embodiments of this application. In still other embodiments, the conference system further includes a camera device for acquiring image signals. The number of terminal devices in the conference system can be one or more.
[0066] It should be understood that the system architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0067] Figure 2 This is a flowchart illustrating an audio processing method provided in an embodiment of this application. The method is applied to an audio-visual system, which can be any of the audio-visual systems described above, such as a conferencing system. The steps of the method can be executed by a processor within the audio-visual system. Please refer to... Figure 2 The method includes the following steps.
[0068] Step 201: Obtain the initial audio signal, which includes the speaker's audio signal and noise signal.
[0069] In an audio / video system that includes a microphone, the initial audio signal comprises the audio signal captured by that microphone. After capturing the initial audio signal, the microphone sends it to the processor, which then receives the initial audio signal.
[0070] In an audio-visual system comprising multiple microphones, the initial audio signal includes the audio signals collected by each of these microphones. Each microphone sends its collected audio signal to the processor, which receives these signals to obtain the initial audio signal. The noise signal in the initial audio signal includes the noise signals collected by each of the microphones, and the speaker's audio signal in the initial audio signal includes the speaker's audio signal collected by each of the microphones. Of course, there may be cases where some microphones fail to collect the speaker's audio signal. In such cases, the speaker's audio signal in the initial audio signal includes the speaker's audio signal collected by the first group of microphones, which refers to the microphones among the multiple microphones that collected the speaker's audio signal.
[0071] It should be understood that the initial audio signal is an audio signal acquired in real time, such as a conference audio signal acquired in real time.
[0072] In this embodiment, the number of speakers can be one or more, and any person in the scene can become a speaker at any time and can also become a non-speaker at any time. For example, if there is originally one speaker in the scene, a new speaker can be added at the first moment, making a total of two speakers. At the second moment, another speaker can be added, making a total of three speakers. At the third moment, one of the speakers can leave. Based on this, the initial audio signal includes the audio signals of one or more speakers.
[0073] Step 202: Determine the speaker's location information.
[0074] It should be understood that the overall concept of this application embodiment is to dynamically determine the audio-visual range based on the speaker's real-time location. Therefore, the processor needs to determine the speaker's location information, which is the speaker's real-time location information.
[0075] There are several ways to determine the speaker's location information; the following section introduces two of them.
[0076] The first implementation method determines the speaker's location information through facial recognition. That is, the processor acquires an image signal and determines the speaker's location based on this signal, where the image signal includes the speaker's image information. This implementation method offers high accuracy and reliability. It should be understood that this image signal is acquired in real-time. The image signal can include a single frame or multiple consecutive frames.
[0077] There are several ways for the processor to determine the speaker's location information based on the image signal. For example, it can use a deep neural network to determine the speaker's location information based on the image signal, or it can determine the speaker's location information by comparing target features based on the image signal.
[0078] The processor can first detect the speaker in the image signal, and then determine the speaker's location information based on the image signal. For example, the processor first detects whether a specified behavior exists in the image signal (which can be detected using deep neural networks or other methods). If the specified behavior is detected, the person performing the specified behavior is identified as the speaker, and thus the speaker's location information is determined. The specified behavior can be actions such as raising a hand, waving, or standing up. Taking raising a hand as an example, if someone raises their hand, the processor can detect the hand-raising behavior and identify the person raising their hand as the speaker. Alternatively, if the processor stores the speaker's facial features, it can detect whether the speaker's facial image information exists in the image signal through target feature comparison. If facial image information is detected, the speaker is identified and their location information is determined.
[0079] There are many ways for the processor to determine the speaker's location information based on the image signal. For example, the processor can determine the speaker's location information based on the detected position of the speaker in the image signal and the actual coordinates of each region in the image signal. Another example is that the processor can determine the speaker's distance and orientation angle relative to the camera based on the detected position and size of the speaker in the image signal and the camera's position coordinates, and then determine the speaker's location information based on this distance and orientation angle.
[0080] In this embodiment of the application, after detecting the speaker, the processor can use a target following method to continuously follow the speaker, that is, to determine the speaker's location information in real time through target following.
[0081] The second method involves determining the speaker's location information through sound source localization.
[0082] For example, the speaker's azimuth angle can be determined using a single microphone; this azimuth angle represents the speaker's direction relative to that microphone. Alternatively, the speaker's azimuth angle can be determined using at least two microphones, resulting in at least two azimuth angles. These at least two azimuth angles represent the speaker's direction relative to their respective microphones, and the overlapping position of these at least two directions is then determined as the speaker's position coordinates. Specific implementation methods can be found in related technologies, which will not be detailed here.
[0083] In this embodiment, the speaker's location information includes the speaker's orientation angle and / or position coordinates. Both the orientation angle and position coordinates indicate the speaker's direction relative to the microphone, and the position coordinates also indicate the distance between the speaker and the microphone.
[0084] In the implementation of determining the speaker's location information through human face detection, the processor can determine the speaker's direction (and / or distance) relative to the microphone based on the camera's position, the microphone's position, and the aforementioned image signal, thereby obtaining the aforementioned direction angle (and / or position coordinates).
[0085] If the scene includes multiple microphones, then in one possible implementation, the processor can determine the direction (and / or distance) of the speaker relative to each of the multiple microphones based on the image signal, and obtain the direction angle (and / or position coordinates) corresponding to each of the multiple microphones.
[0086] In the method of determining the speaker's location information through sound source localization, the processor can determine the speaker's direction (and / or distance) relative to each of the at least two microphones based on the audio signals collected by at least two microphones, and obtain the azimuth angle (and / or position coordinates) corresponding to the at least two microphones respectively. Alternatively, the processor can also determine the speaker's direction relative to a single microphone based on the audio signal collected by one microphone, and obtain a azimuth angle.
[0087] In some embodiments, the speaker's position information may also include pitch angle, distance, etc. That is, in addition to determining the speaker's azimuth angle and / or position coordinates, the processor can also determine the speaker's pitch angle and distance relative to the microphone. Specific implementation methods can be found in related technologies, which will not be detailed here.
[0088] Step 203: Determine the audio-visual range based on the speaker's location information.
[0089] In other words, the processor dynamically determines the audio-visual range based on the speaker's real-time location information.
[0090] In this embodiment, the processor can update the audio veil range when the speaker's position changes, or it can determine the audio veil range in real time. These two implementation methods will be described below.
[0091] In the first implementation, the processor updates the audio gaiter range when the speaker's position changes.
[0092] That is, the speaker's location information in step 202 above includes the location information at the first moment and the location information at the second moment, with the first moment preceding the second moment. When the processor determines that the speaker's location has changed based on the location information at the first and second moments, it determines the audio-visual censorship range based on the location information at the second moment.
[0093] In this embodiment, the first moment and the second moment can be two consecutive frames. These two consecutive frames can be two image signals continuously acquired by the camera in a face detection method, or two audio signals continuously acquired by the microphone in a sound source localization method. Alternatively, the first moment and the second moment can be spaced apart by a sampling period. Within one sampling period, the camera can acquire one or more image signals, and the microphone can acquire one or more audio signals. This embodiment does not limit the scope of the application.
[0094] The second implementation method involves the processor determining the audio veil range in real time.
[0095] In other words, the speaker's location information in step 202 above includes the location information of the current frame. The processor determines the audio-visual range based on the location information of the current frame.
[0096] In the human face detection method, the current frame corresponds to a frame of image signal currently captured by the camera. In the sound source localization method, the current frame corresponds to a frame of audio signal currently captured by the microphone.
[0097] As mentioned above, the speaker's location information includes the speaker's orientation angle and / or position coordinates. That is, the specific information included in the speaker's location information varies depending on the positioning granularity, and the shape of the audio-visual area that the processor can determine can differ depending on the specific information included in the speaker's location information. This will be discussed in the following sections.
[0098] In the first scenario, the speaker's location information includes the speaker's azimuth angle and / or position coordinates, which indicate the speaker's direction relative to the target microphone, and the audio gait range includes the area pointed to by that direction.
[0099] The target microphone is the microphone that captures the speaker's audio signal; that is, the microphone corresponding to the audio signal that will ultimately be played. It should be understood that if the scene includes one microphone, that microphone is the target microphone. If the scene includes multiple microphones, the target microphone is one of those microphones; that is, the audio signal of the speaker captured by one of those microphones will ultimately be played. During step 202, the processor can determine the speaker's position relative to the target microphone, but not the speaker's position relative to the other microphones. Alternatively, it can determine the speaker's position relative to each of all microphones, although the speaker's position relative to the other microphones may not be used in subsequent processing steps.
[0100] In one possible implementation, see Figure 3 The audio censorship range determined by the processor includes the area between the first ray and the second ray, which intersect at the target microphone. The speaker's direction relative to the target microphone is located between the first and second rays. In some embodiments, the audio censorship range corresponding to this audio censorship range can be referred to as a fan-shaped audio censorship. That is, Figure 3 This is a schematic diagram of a fan-shaped sound screen provided in an embodiment of this application. In some embodiments, the first ray and the second ray can be rays from the target microphone pointing to the left and right sides of the speaker, respectively.
[0101] In this embodiment of the application, when the speaker's position information includes the speaker's direction angle and / or position coordinates, one implementation of the processor determining the audio-visual screen range based on the speaker's position information is as follows: Based on the speaker's position information and a reference angle, a first ray and a second ray are determined; based on the first ray and the second ray, the audio-visual screen range is determined. Wherein, the angle between the first ray and the direction, and the angle between the second ray and the direction, are both equal to the reference angle; or, the angle between the first ray and the direction, and the angle between the second ray and the direction, are both equal to half of the reference angle; or, the reference angle includes a first reference angle and a second reference angle, the angle between the first ray and the direction is equal to the first reference angle, and the angle between the second ray and the direction is equal to the second reference angle.
[0102] The reference angle can be a preset angle or an angle determined based on the distance between the speaker and the target microphone. For example, if the speaker is far from the target microphone, the processor can determine the reference angle as a first value; if the speaker is close to the target microphone, the processor can determine the reference angle as a second value, where the first value is greater than the second value. In one possible implementation, the processor can determine the reference angle based on the distance between the speaker and the target microphone and the mapping relationship between distance and reference angle. This mapping relationship can be a mapping table or a mapping function, or it can be an implicit mapping relationship represented by a neural network model; this embodiment does not limit this.
[0103] When the reference angle includes a first reference angle and a second reference angle, the first reference angle can be the same as the second reference angle, or they can be different. Both the first and second reference angles can be preset angles, or they can both be angles determined based on the distance between the speaker and the target microphone. The specific implementation method is similar to the implementation method of the processor determining the reference angle based on the distance between the speaker and the target microphone in the previous paragraph, and will not be described in detail here.
[0104] For example, the reference angle can be 15 degrees, 20 degrees, or 30 degrees, etc. The first reference angle and / or the second reference angle can be 10 degrees, 12 degrees, or 15 degrees, etc.
[0105] It should be understood that the sound screen range is a three-dimensional range. A specific implementation of the processor determining the sound screen range based on the first ray and the second ray can be as follows: the sound screen range is determined based on the first ray, the second ray, and the vertical height. The height of the sound screen range is equal to the vertical height. The range of the sound screen range in the vertical direction can be the range from the ground to the vertical height above the ground. The range of the sound screen range in the horizontal direction can be the range between the orthographic projection of the first ray on the horizontal plane and the orthographic projection of the second ray on the horizontal plane.
[0106] The vertical height can be a preset height, which can be set according to the actual scenario, such as the ceiling height of the conference room. For example, the vertical height can be 2.5 meters, 3 meters, or 5 meters, etc.
[0107] In the second scenario, the speaker's location information includes their coordinates, and the processor determines the audio-visual censorship range to include the area centered on those coordinates. The audio-visual censorship range corresponding to this censorship range can be called a fixed-point audio-visual censorship.
[0108] In one implementation, see Figure 4The audio screen range includes an area centered on the speaker's position coordinates; that is, the audio screen range corresponds to a circular audio screen, which can be called a circular fixed-point audio screen. In another implementation, the audio screen range includes a polygonal area centered on the speaker's position coordinates; that is, the audio screen range corresponds to a polygonal audio screen, which can be a square, rectangular, or pentagonal area, etc. Besides circular and polygonal audio screens, other shapes of audio screens are also possible; this application does not limit the shape of the audio screen. It should be understood that the circle and polygon here can refer to the orthographic projection of the audio screen range onto a horizontal plane; the audio screen range itself is a three-dimensional area.
[0109] Taking a circular sound screen as an example, the processor can first determine a circular area with the speaker's position coordinates as the center and a reference distance as the radius, and then determine the sound screen range based on this circular area. In one possible implementation, the processor can determine the sound screen range based on the circular area and the vertical height. The height of the sound screen range is equal to the vertical height. The vertical range of the sound screen range can be from the ground to the vertical height above the ground, and the horizontal range of the sound screen range can be the range between the orthographic projection of the circular area on the horizontal plane and the orthographic projection of the circular area on the horizontal plane. The circular area can be a horizontal region. Thus, the circular sound screen is a cylindrical sound screen.
[0110] The reference distance is a preset distance, which can be 1 meter, 2 meters, or 3 meters, etc. The setting of the reference distance can be determined according to the size of the venue and the precision of the audio-visual technology. The vertical height here can be the same as the vertical height mentioned above. Its possible values and setting basis can be found in the relevant introduction above, and will not be repeated here.
[0111] Alternatively, in another possible implementation, the processor can define the audio-visual area as a spherical region centered on the speaker's location coordinates and with a reference distance as its radius.
[0112] In this embodiment, the number of speakers is M, where M is a positive integer. When M is greater than 1, the audio-visual range determined by the processor includes the audio-visual ranges corresponding to each of the M speakers. For any one of the M speakers, the processor can determine the corresponding audio-visual range according to the method described above, which will not be repeated here.
[0113] In one possible implementation, the shape and / or size of the audio-visual screen corresponding to these M speakers may be the same or different, and this application embodiment does not limit this. Figure 5 and Figure 6 These are schematic diagrams of two types of multi-person audio-visual subtitles provided in the embodiments of this application. Figure 5and Figure 6 There are 3 main speakers. Figure 5 The audio screens for the three main speakers are all circular. Figure 6 The audio screens for the three main speakers are all fan-shaped.
[0114] Step 204: Based on the audio censorship range, eliminate noise signals in the initial audio signal.
[0115] In a microphone scenario, the initial audio signal includes the speaker's audio signal and noise signal collected by the microphone. Based on the audio gamut range corresponding to the microphone (which may be called the target microphone), the processor eliminates the noise signal collected by the microphone to obtain the target audio signal, which includes the speaker's audio signal.
[0116] In scenarios with multiple microphones, the initial audio signal includes the audio signals collected by each of those microphones. The spectral range is the spectral range corresponding to the target microphone, which is one of the multiple microphones. The processor can, based on this spectral range, remove noise signals from the audio signal collected by the target microphone to obtain the target audio signal. The target audio signal includes the speaker's audio signal collected by the target microphone. It should be understood that in this scenario, only the speaker's audio signal collected by the target microphone is played; the audio signals collected by other microphones do not need to be processed and are not played.
[0117] As can be seen from the above, there can be one or more speakers, so the target audio signal can also include the audio signals of one or more speakers.
[0118] In this embodiment, there are many specific ways for the processor to eliminate noise signals collected by the target microphone based on the sound censorship range. For example, it can eliminate all sound signals whose sound sources are outside the sound censorship range, including human voice signals whose sound sources are outside the sound censorship range, based on sound source localization. It can also eliminate white noise (also known as background noise or environmental noise) collected by the target microphone. Of course, the processor can also eliminate the above-mentioned noise signals in other ways, and this embodiment does not limit the specific implementation method.
[0119] In addition to eliminating noise from the initial audio signal, the processor can also enhance the audio signal of the speaker captured by the target microphone to further improve the speaker's audio quality. There are various methods for audio enhancement, such as software processing through deep neural networks or hardware methods like amplifiers; this application does not limit the specific methods used.
[0120] Considering the scenario involving multiple microphones, the distances between these microphones and the speaker vary. Some microphones may be farther away, while others may be closer. Generally speaking, the closer the microphone is to the speaker, the higher the quality of the audio signal it captures. Therefore, the target microphone can be the one closest to the speaker among the multiple microphones, thus maximizing the quality of the target audio signal. Specifically, when there are multiple speakers, the target microphone can be the one with the closest average distance to all speakers.
[0121] Alternatively, in some embodiments, the target microphone may be the microphone corresponding to the median distance, where the median distance refers to the median distance between the plurality of microphones and the plurality of speakers. Alternatively, in still other embodiments, the target microphone may be selected according to other rules, which are not limited in this application.
[0122] As described above, in scenarios with multiple microphones, these microphones can be distributed across multiple terminal devices, with each terminal device containing at least one microphone. In the case where the processor is not one of these terminal devices—for example, if the processor is a processor in a control device or server—after determining the speaker's location information, the processor can synchronize the speaker's location information with the multiple terminal devices, instructing each terminal device to determine its distance from the speaker. After each terminal device determines its distance from the speaker based on the speaker's location information, it sends distance information to the processor. This distance information represents the distance between the respective terminal device and the speaker. The processor receives the distance information sent by each terminal device and, based on this information, determines the target microphone from among the multiple microphones; for example, it identifies the microphone closest to the speaker as the target microphone. Alternatively, after determining the speaker's location information, the processor can determine the distance between the speaker and each of the multiple terminal devices based on the location information of the multiple terminal devices and the speaker's location information, and then determine the target microphone from the multiple microphones based on the distance between the speaker and each of the multiple terminal devices.
[0123] When the processor is the processor of the first terminal device among the plurality of terminal devices, the processor can synchronize the speaker's location information with the other terminal devices, receive distance information sent by the other terminal devices, and then determine the target microphone from the plurality of microphones based on the received distance information and the distance between the first terminal device and the speaker. Of course, in some embodiments, the processor may also store the location information of the plurality of terminal devices, and then determine the target microphone from the plurality of microphones based on the location information of the plurality of terminal devices and the speaker's location information.
[0124] In one possible implementation, the multiple terminal devices can also determine the speaker's location independently, without the aforementioned synchronization operation. For example, each of the multiple terminal devices includes a camera, and each terminal device acquires image signals captured by its own camera, determining the speaker's location information based on the acquired image signals.
[0125] In this embodiment, considering that the speaker may change, for example, a person may become the speaker by raising their hand or standing up, and if that person raises their hand again or sits down, then that person will no longer be the speaker. Based on this, the processor can also detect changes in the speaker, such as determining, based on image signals, that a person has changed from the speaker to a non-speaker, and thus turn off the audio / video screen corresponding to that person.
[0126] Next, please combine... Figure 7 The modules shown are used to illustrate the audio processing method provided in the embodiments of this application.
[0127] Figure 7 This is a schematic diagram of an audio processing device provided in an embodiment of this application. See also... Figure 7 The audio processing device includes a video acquisition module, a video processing module, a video output module, an audio acquisition module, an audio processing module, an audio output module, a control module, and a network transmission module. The video processing module includes sub-modules for video encoding / decoding, target detection, and target tracking, while the audio processing module includes sub-modules for audio encoding / decoding, sound source localization, and intelligent audio veiling. The steps of the aforementioned audio processing method can be executed by this audio processing device.
[0128] The video-related modules, control module, and network transmission module are all optional. The video acquisition module mainly includes one or more cameras to acquire image signals (which can be image signals from video signals) as input to the subsequent video processing and output modules. The audio acquisition module mainly includes one or more microphones to acquire initial audio signals as input to the subsequent audio processing module. The sound source localization module, a submodule of the audio processing module, is used in the intelligent audio censorship algorithm to determine the location of the sound source in the initial audio signal and whether the sound source is within the audio censorship range. The target detection module and target tracking module, two submodules of the video processing module, are mainly used for target detection (e.g., face detection) and target tracking based on the image signals acquired by the cameras to determine the speaker's location information in real time. The speaker's location information can be used by the intelligent audio censorship submodule in the audio processing module to adjust the audio censorship range. The control module can be used to control the audio and video processing modules. The network transmission module can be used to transmit the audio signals output by the audio processing module and the image signals output by the video processing module over the network, and can also be used to transmit audio and image signals from the network to the audio and video processing modules respectively.
[0129] In summary, the embodiments of this application can dynamically determine the audio censorship range based on the speaker's real-time location information, allowing the audio censorship range to change with the speaker's position, thereby ensuring that the speaker is always within the audio censorship range and improving sound pickup quality. In one specific implementation, target detection and target tracking of the speaker can be performed based on image signals, and then the audio censorship range can be dynamically updated based on the speaker's real-time location information.
[0130] Figure 8 This is a schematic diagram of another audio processing device provided in an embodiment of this application. This audio processing device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be... Figure 1 The computer equipment shown. See also Figure 8 The audio processing device includes: an acquisition module 801, a first determination module 802, a second determination module 803, and an audio processing module 804.
[0131] The acquisition module 801 is used to acquire the initial audio signal, which includes the speaker's audio signal and noise signal;
[0132] The first determining module 802 is used to determine the location information of the speaker;
[0133] The second determining module 803 is used to determine the audio-visual range based on the speaker's location information;
[0134] The audio processing module 804 is used to eliminate noise signals in the initial audio signal based on the audio censoring range.
[0135] In one possible implementation, the speaker's location information includes location information at a first moment and location information at a second moment, with the first moment preceding the second moment; the second determining module 803 includes:
[0136] The first determining submodule is used to determine the audio-visual range based on the location information at the second moment, in the case where the speaker's position has changed based on the location information at the first moment and the location information at the second moment.
[0137] In one possible implementation, the speaker's location information includes the location information of the current frame; the second determining module 803 includes:
[0138] The second determination submodule is used to determine the audio-visual range based on the position information of the current frame.
[0139] In one possible implementation, the speaker's location information includes the speaker's azimuth angle and / or position coordinates, which indicate the speaker's direction relative to the target microphone, the audio gait range includes the area pointed to by that direction, and the target microphone is a microphone that collects the speaker's audio signal.
[0140] In one possible implementation, the sound curtain range includes the area between a first ray and a second ray, which intersect at the target microphone, with the aforementioned direction located between the first ray and the second ray.
[0141] In one possible implementation, the speaker's location information includes the speaker's location coordinates, and the audio-visual range includes the area centered on those location coordinates.
[0142] In one possible implementation, the audio-visual area includes the region centered at the location coordinates.
[0143] In one possible implementation, the first determining module 802 includes:
[0144] The acquisition submodule is used to acquire image signals, which include image information of the speaker;
[0145] The third determination submodule is used to determine the speaker's location information based on the image signal.
[0146] In one possible implementation, the initial audio signal includes audio signals collected by multiple microphones, the soundstage range is the soundstage range corresponding to the target microphone, and the target microphone is one of the multiple microphones; the audio processing module 804 includes:
[0147] The audio processing submodule is used to remove noise signals from the audio signal captured by the target microphone based on the sound field range.
[0148] In one possible implementation, the target microphone is the microphone closest to the speaker among the plurality of microphones.
[0149] In one possible implementation, the plurality of microphones are distributed across multiple terminal devices, each terminal device including at least one microphone, and the audio processing device further includes:
[0150] The synchronization module is used to synchronize the speaker's location information with the multiple terminal devices;
[0151] The receiving module is used to receive distance information sent by the multiple terminal devices respectively, which represents the distance between the corresponding terminal device and the speaker;
[0152] The third determining module is used to determine the target microphone from the multiple microphones based on the distance information sent by the multiple terminal devices.
[0153] In one possible implementation, the plurality of microphones are distributed across multiple terminal devices, each terminal device including at least one microphone, and the audio processing device further includes:
[0154] The fourth determining module is used to determine the distance between the speaker and each of the multiple terminal devices based on the location information of the multiple terminal devices and the location information of the speaker;
[0155] The fifth determination module is used to determine the target microphone from the multiple microphones based on the distance between the speaker and the multiple terminal devices.
[0156] In one possible implementation, the number of speakers is M, where M is a positive integer; and when M is greater than 1, the audio-visual scope includes the audio-visual scope corresponding to each of the M speakers.
[0157] This application embodiment can dynamically determine the audio censorship range based on the speaker's real-time location information, allowing the audio censorship range to change with the speaker's position, thereby ensuring that the speaker is always within the audio censorship range and improving sound pickup quality. In one specific implementation, target detection and target tracking of the speaker can be performed based on image signals, and then the audio censorship range can be dynamically updated based on the speaker's real-time location information.
[0158] It should be noted that the audio processing device provided in the above embodiments is only illustrated by the division of the above functional modules when processing audio signals. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device and the audio processing method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0159] This application also provides another audio processing apparatus, which includes a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to implement the audio processing method shown in the above method embodiments.
[0160] In one possible implementation, the audio processing device is included in the control device, any terminal device, or server in the above embodiments. For example, the audio processing device is included in the control device of the conference system, and the control device can be a central control device.
[0161] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the steps of the audio processing method shown in the above-described method embodiments.
[0162] This application also provides a computer program product containing instructions that, when run on a computer, causes the computer to perform the steps of the audio processing method shown in the above-described method embodiments. Alternatively, this application also provides a computer program that, when run on a computer, causes the computer to perform the steps of the audio processing method shown in the above-described method embodiments.
[0163] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.
[0164] It should be understood that "at least one" as mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., are not necessarily different.
[0165] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the audio signals involved in the embodiments of this application were all obtained under full authorization.
[0166] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An audio processing method, characterized in that, The method includes: Acquire an initial audio signal, which includes the speaker's audio signal and a noise signal; Determine the location information of the speaker; Based on the speaker's location information, the audio-visual area is determined; Based on the audio censorship range, noise signals in the initial audio signal are eliminated.
2. The method as described in claim 1, characterized in that, The speaker's location information includes location information at a first moment and location information at a second moment, with the first moment occurring before the second moment. Determining the audio-visual range based on the speaker's location information includes: If the speaker's position changes based on the position information at the first moment and the position information at the second moment, the audio-visual range is determined based on the position information at the second moment.
3. The method as described in claim 1, characterized in that, The speaker's location information includes the location information of the current frame; Determining the audio-visual range based on the speaker's location information includes: The audio-visual range is determined based on the position information of the current frame.
4. The method according to any one of claims 1-3, characterized in that, The speaker's location information includes the speaker's azimuth angle and / or position coordinates, the azimuth angle and / or position coordinates indicating the speaker's direction relative to the target microphone, the audio gait range including the area indicated by the direction, and the target microphone being a microphone that collects the speaker's audio signals.
5. The method as described in claim 4, characterized in that, The sound curtain range includes the area between the first ray and the second ray, the first ray and the second ray intersecting at the target microphone, and the direction being located between the first ray and the second ray.
6. The method according to any one of claims 1-3, characterized in that, The speaker's location information includes the speaker's location coordinates, and the audio-visual range includes the area centered on the location coordinates.
7. The method as described in claim 6, characterized in that, The audio-visual area includes the region centered at the stated location coordinates.
8. The method according to any one of claims 1-7, characterized in that, Determining the location information of the speaker includes: Acquire an image signal, the image signal including the image information of the speaker; Based on the image signal, the location information of the speaker is determined.
9. The method according to any one of claims 1-8, characterized in that, The initial audio signal includes audio signals collected by multiple microphones respectively, the sound field range is the sound field range corresponding to the target microphone, and the target microphone is one of the multiple microphones; The step of eliminating noise signals in the initial audio signal based on the audio censorship range includes: Based on the sound gaiter range, noise signals in the audio signal collected by the target microphone are eliminated.
10. The method as described in claim 9, characterized in that, The target microphone is the microphone that is closest to the speaker among the plurality of microphones.
11. The method as described in claim 9 or 10, characterized in that, The plurality of microphones are distributed across multiple terminal devices, and each terminal device includes at least one microphone. The method further includes: Synchronize the speaker's location information with the multiple terminal devices; The system receives distance information sent by the plurality of terminal devices, wherein the distance information represents the distance between the corresponding terminal device and the speaker. The target microphone is determined from the plurality of microphones based on the distance information sent by the plurality of terminal devices.
12. The method as described in claim 9 or 10, characterized in that, The plurality of microphones are distributed across multiple terminal devices, and each terminal device includes at least one microphone. The method further includes: Based on the location information of the multiple terminal devices and the location information of the speaker, the distance between the speaker and the multiple terminal devices is determined. The target microphone is determined from the plurality of microphones based on the distance between the speaker and the plurality of terminal devices.
13. The method according to any one of claims 1-12, characterized in that, The number of speakers is M, where M is a positive integer; wherein, when M is greater than 1, the audio-visual range includes the audio-visual ranges corresponding to the M speakers respectively.
14. An audio processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire an initial audio signal, which includes the speaker's audio signal and a noise signal; The first determining module is used to determine the location information of the speaker; The second determining module is used to determine the audio-visual range based on the speaker's location information; An audio processing module is used to eliminate noise signals in the initial audio signal based on the audio censorship range.
15. An audio processing apparatus, characterized in that, The device includes a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method according to any one of claims 1-13.
16. A conference system, characterized in that, The conference system includes terminal equipment and control equipment; The terminal device is used to collect initial audio signals; The control device is used to perform the steps of the method according to any one of claims 1-13.
17. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1-13.
18. A computer program product, characterized in that, The computer program product stores computer instructions, which, when executed by a processor, implement the method described in any one of claims 1-13.