Audio processing methods, electronic devices and storage media
By identifying the location of the target sound source through the image acquisition component, generating an acoustic image, and determining the target beam and volume gain, the problem of inconsistent sound in remote meetings is solved, thereby improving audio quality and meeting efficiency.
Patent Information
- Application Number
- CN202211250411.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-10-12
AI Technical Summary
In remote meetings, the difference in volume and distance between each participant and the sensor can lead to significant differences in sound quality, which can negatively impact the meeting's quality.
The image acquisition component identifies the location information of the target sound source, generates an audio image, determines the relative positional relationship between the target sound source and the audio acquisition component, sets the target beam and volume gain, and adjusts the volume of the audio data.
It improves the quality of conference audio data, ensures consistent volume across all target sound sources, reduces noise interference, and enhances conference efficiency and equipment reliability.
Smart Images

Figure CN115633290B_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an audio processing method, electronic device, and storage medium, belonging to the field of audio processing technology. [Background Technology]
[0002] When conducting remote meetings, a conference terminal is needed to connect different meeting rooms. These terminals typically have sensors for collecting sound. However, noise, echoes, reverberation, and other interference factors inevitably exist within the meeting room, severely affecting the quality of sound transmitted to another meeting room and disrupting the normal progress of the meeting.
[0003] Existing conference terminals use sensors to set a fixed beam facing the participants, collecting only the sound within the beam to reduce interference from sounds outside the beam, thereby mitigating the impact of interference factors in the conference room.
[0004] However, because each participant's voice volume is different, and the distance between each participant and the sensor is also different, this can lead to a significant difference in the voices of the participants. [Summary of the Invention]
[0005] This application provides an audio processing method, electronic device, and storage medium that addresses the problem of significant differences in sound between participants due to variations in volume and distance between each participant and the sensor. The technical solution provided in this application is as follows:
[0006] Firstly, an audio processing method is provided, the method comprising:
[0007] Image recognition is performed on the images acquired by the image acquisition component within the conference scene to obtain the first location information of the target sound source;
[0008] A sound image is generated based on the audio signal acquired by the audio acquisition component. The sound image is used to indicate the relative positional relationship between each sound source and the audio acquisition component in the conference scene.
[0009] The first location information is used to determine the target relative positional relationship between the target sound source and the audio acquisition component in various relative positional relationships;
[0010] The target beam and volume gain of the target sound source are determined based on the relative positional relationship of the targets, wherein the volume gain is positively correlated with the distance indicated by the relative positional relationship of the targets;
[0011] The volume gain is used to adjust the audio volume of the audio data acquired according to the target beam.
[0012] Optionally, the first location information includes the target orientation and the target distance, wherein the target orientation includes the horizontal angular orientation and the vertical angular orientation;
[0013] The step of performing image recognition on images acquired by the image acquisition component within the conference scene to obtain the first location information of the target sound source includes:
[0014] Face detection is performed on the image to obtain the face pixel height and the coordinates of the face center point;
[0015] The screen pixel size, screen center point coordinates, horizontal field of view, vertical field of view, and true face height of the image acquisition component are obtained. The screen pixel size includes both horizontal and vertical dimensions.
[0016] Obtain the horizontal and vertical distances between the coordinates of the face center point and the coordinates of the screen center point;
[0017] The horizontal angular orientation between the face center point coordinates and the image acquisition component is obtained using the horizontal dimension, the horizontal distance, and the horizontal field of view.
[0018] The vertical angular orientation between the face center point coordinates and the image acquisition component is obtained using the vertical dimension, the vertical distance, and the vertical field of view.
[0019] The focal length is obtained using the vertical dimension and the vertical field of view.
[0020] The target distance is obtained using the focal length, the actual height of the face, and the pixel height of the face.
[0021] Optionally, the acoustic image includes audio peaks of each sound source, the positions of which are used to indicate the relative positional relationships; determining the target relative positional relationship between the target sound source and the audio acquisition component using the first positional information in the relative positional relationships includes:
[0022] Obtain the second location information corresponding to the first location information in the acoustic image;
[0023] Among the various audio peaks in the sound image, identify the false audio peaks whose relative positional relationship does not match the second positional information;
[0024] The false audio peaks are removed from the audio peaks in the audio image to obtain the target relative position relationship indicated by each of the removed audio peaks.
[0025] Optionally, the second location information includes a first distance and a first orientation between the audio acquisition component and the target sound source, and the relative positional relationships include a second distance and a second orientation;
[0026] The step of identifying false audio peaks whose relative positional relationship does not match the second positional information among the various audio peaks in the acoustic image includes:
[0027] In the various audio peaks of the acoustic image, the distance difference between the second distance and the first distance is obtained.
[0028] The false audio peaks are identified as those where the second orientation does not match the first orientation, and / or the distance difference is greater than a preset distance threshold.
[0029] Optionally, determining false audio peaks whose relative positional relationships do not match the second positional information among the various audio peaks in the acoustic image includes:
[0030] If the number of audio peaks in the sound image is greater than or equal to the number of target sound sources, false audio peaks whose relative positional relationship does not match the second positional information are identified among the audio peaks in the sound image.
[0031] Optionally, determining the volume gain of the target sound source based on the relative positional relationship of the targets includes:
[0032] The intermediate distance is determined using the maximum and minimum distances indicated by the relative positional relationship of the targets;
[0033] Obtain the standard volume gain corresponding to the intermediate distance;
[0034] Based on the difference between each target sound source and the intermediate distance, the gain difference between the volume gain of each target sound source and the standard volume gain is determined;
[0035] The volume gain of each target sound source is determined based on the standard volume gain and the gain difference.
[0036] Optionally, determining the volume gain of the target sound source based on the relative positional relationship of the targets includes:
[0037] Based on the relative positional relationship of the targets, a first gain value of the target sound source is determined, wherein the first gain value is positively correlated with the distance indicated by the relative positional relationship of the targets;
[0038] Determine the volume difference between the audio volume obtained by adjusting the audio data according to the first gain value and the preset standard volume;
[0039] The volume difference is used to determine the second gain value of the target sound source;
[0040] The volume gain of the target sound source is determined based on the first gain value and the second gain value.
[0041] Optionally, the method includes:
[0042] Determine whether the target sound source meets preset volume adjustment conditions; the volume adjustment conditions include at least one of the following: the target sound source is in a speaking state; the target sound source is a speaker; the target sound source is a participant interacting with the speaker;
[0043] If the target sound source meets the volume adjustment conditions, the step of adjusting the volume of the audio data acquired according to the target beam using the volume gain is triggered.
[0044] Optionally, the method further includes:
[0045] If the target sound source does not meet the volume adjustment conditions, and the volume of the audio data is greater than or equal to a preset volume threshold, then the audio volume of the audio data is suppressed.
[0046] In a second aspect, an electronic device is provided, the electronic device including a processor and a memory connected to the processor, the memory storing a program, wherein the processor executes the program to implement the audio processing method provided in the first aspect.
[0047] Thirdly, a computer-readable storage medium is provided, wherein a program is stored therein, which, when executed by a processor, is used to implement the audio processing method provided in the first aspect.
[0048] The beneficial effects of this application include at least the following: obtaining the first location information of the target sound source by performing image recognition on images acquired by the image acquisition component within the conference scene; generating an acoustic image based on the audio signal acquired by the audio acquisition component, the acoustic image being used to indicate the relative positional relationships between each sound source and the audio acquisition component within the conference scene; determining the target relative positional relationship between the target sound source and the audio acquisition component using the first location information in each relative positional relationship; determining the target beam and volume gain of the target sound source based on the target relative positional relationship; adjusting the audio volume of the audio data acquired according to the target beam using the volume gain; solving the problem of significant differences in sound between participants due to variations in the volume of each participant's voice and the distance between each participant and the sensor; and since the acoustic image is used to indicate the conference scene... The relative positions of all sound sources within the audio image are determined, including the relative positions of sound sources affected by noise, echoes, reverberation, and other interference factors. Using an image allows for the identification of the target sound source and the acquisition of its initial positional information. Therefore, by using this initial positional information to determine the target relative position within the various relative positions, the relative positions of noise, echoes, reverberation, and other interference factors can be eliminated in the audio image, thus obtaining the target relative position of the sound source. Since the relative positional information obtained from the audio image is relatively accurate, the precise positional information of the target sound source can be obtained. Target beamforming and volume gain can then be set for the target sound source, ensuring that the volume differences between target sound sources are not too large. This improves the quality of audio data within the conference and ensures that the audio volume of each target sound source is consistent.
[0049] Meanwhile, by combining the first position information and various relative positional relationships to determine the target relative positional relationship between the target sound source and the audio acquisition component, the problem of the target beam deviating from the target sound source due to inaccurate first position information caused by image distortion can be avoided; it can also avoid the problem of the target beam pointing to noise due to the formation of audio peaks in the sound image caused by noise; by combining the first position information and various relative positional relationships to determine the target relative positional relationship corresponding to the target sound source, the target beam pointing to the target sound source can be accurately set, and the audio volume of the audio data of the target sound source acquired in the target beam can be adjusted by adjusting the volume gain, so that the volume of each target sound source tends to be consistent and conforms to the user's listening habits, thereby improving the quality of conference audio.
[0050] At the same time, by using volume gain adjustment to adjust the audio gain of the audio data of the target sound source collected within the target beam, the problem of participants being unable to hear clearly due to excessively loud or soft volume can be avoided. Therefore, the number of times participants have to repeat themselves because they cannot hear clearly can be reduced, thus shortening meeting time and improving meeting efficiency.
[0051] In addition, by deleting false audio peaks that do not match the first position information from the audio image, the relative positional relationship of the target indicated by each audio peak after deletion can be obtained. This can improve the accuracy of the obtained relative positional relationship of the target, thereby improving the accuracy of the target beam pointing to the target sound source and improving the effectiveness of volume gain adjustment of the target sound source volume, thereby reducing the interference of noise and other interference factors in the meeting room on the meeting.
[0052] In addition, when the number of audio peaks is greater than or equal to the number of target sound sources, identifying false audio peaks whose relative positional relationship does not match the first positional information among the various audio peaks in the audio image can avoid errors in obtaining the target relative positional relationship of the target sound sources when the audio components are damaged or not all target sound sources have entered the conference scene, thereby improving the reliability of electronic devices.
[0053] In addition, by acquiring the first location information of the target sound source of the target type and identifying false audio peaks whose relative positional relationship does not match the first location information of the target sound source of the target type, different target sound sources can be flexibly determined according to different meeting scenarios, which can improve the flexibility of electronic devices.
[0054] In addition, by determining the volume gain of each target sound source based on the standard volume gain corresponding to the intermediate distance, the volume of the target sound sources can be made more consistent, which conforms to the user's listening habits. At the same time, it can also reduce the preset correspondence between distance and volume gain, saving storage resources of electronic devices.
[0055] In addition, by determining the second gain value based on the preset standard volume on the basis of the first gain value, the corresponding volume gain can be adjusted in a timely manner when the volume of the target sound source changes, and the volume change can be smoothly transitioned, thereby improving the accuracy of volume adjustment of electronic devices.
[0056] In addition, when the preset volume adjustment conditions are met, using volume gain adjustment to adjust the volume of the audio data collected according to the target beam can ensure that the volume of the target sound source that meets the volume adjustment conditions is highlighted, thereby improving the intelligence of volume adjustment of electronic devices.
[0057] In addition, suppressing the audio volume of the target sound source's audio data when the preset volume adjustment conditions are not met can prevent the target sound source from becoming too loud and interfering with the normal progress of the meeting, thus improving the efficiency of the meeting.
[0058] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. [Attached Image Description]
[0059] Figure 1 This is a schematic diagram of an audio processing method system provided in one embodiment of this application;
[0060] Figure 2 This is a flowchart of an audio processing method provided in one embodiment of this application;
[0061] Figure 3 This is a schematic diagram of an image provided in one embodiment of this application;
[0062] Figure 4 This is a schematic diagram of a target orientation calculation method according to an embodiment of this application;
[0063] Figure 5 This is a schematic diagram of a target distance calculation method according to an embodiment of this application;
[0064] Figure 6 This is a schematic diagram of an audio-visual image provided in one embodiment of this application;
[0065] Figure 7 This is a schematic diagram illustrating the correspondence between a target beam and a target sound source according to an embodiment of this application;
[0066] Figure 8 This is a schematic diagram illustrating the allocation of volume gain based on intermediate distance according to one embodiment of this application;
[0067] Figure 9 This is a flowchart illustrating the specific steps provided in one embodiment of this application;
[0068] Figure 10 This is a block diagram of an apparatus for an audio processing method provided in one embodiment of this application;
[0069] Figure 11 This is a block diagram of an electronic device provided in one embodiment of this application.
Detailed Implementation Methods
[0070] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.
[0071] During remote meetings, the volume of audio data captured by the audio acquisition component can vary significantly due to differences in the speaking volume of each participant and their distance from the audio acquisition unit. This can interfere with the normal progress of the meeting. Traditional solutions include, but are not limited to, at least one of the following:
[0072] The first method involves having participants actively adjust their speaking volume or their distance from the audio acquisition component at one end of the audio acquisition unit to change the volume of the audio data.
[0073] However, participants are unable to properly control the volume and distance they actively adjust, which can cause the volume to change from too high to too low, or from too low to too high, or even cause sound distortion due to cutoff when amplifying the sound.
[0074] The second method involves actively adjusting the output volume of the audio data at the end where the participants' audio is being output.
[0075] While increasing the output volume of audio data, the noise in the audio data will also be amplified; moreover, frequent manual adjustment of the output volume is also quite troublesome.
[0076] The third method involves each participant wearing a corresponding audio capture device, such as a gooseneck microphone.
[0077] Gooseneck microphones require interfaces to connect to terminals. Typically, terminals have a limited number of interfaces, insufficient to accommodate all gooseneck microphones, necessitating external expansion devices. This increases the number of access lines in the meeting setup, raising both the complexity of cabling and costs.
[0078] The fourth method is to use Automatic Gain Control (AGC) technology to automatically adjust the audio volume of the audio data.
[0079] The working principle of AGC (Acoustic Echo Canceller) for volume adjustment is as follows: When a weak signal is input, the linear amplifier circuit operates to ensure the strength of the output signal; when the input signal reaches a certain strength, the compression amplifier circuit is activated, reducing the output amplitude. The greater the similarity between the input and output signals, the better the echo cancellation performance of the Acoustic Echo Chancellor (AEC). Therefore, because AGC increases the difference between the input and output signals, it affects the echo cancellation effect of AEC.
[0080] The fifth method uses a microphone array and beamforming technology to effectively form a beam in the desired direction, picking up only the signal within the beam, thereby improving the signal-to-noise ratio.
[0081] However, due to the changing positions of attendees and the presence of noise, echoes, reverberation, and other interference factors in the meeting environment, the microphone array may be unable to set the beam direction in a timely, accurate, and precise manner according to the attendees' positions.
[0082] Therefore, currently, one or more virtual unidirectional microphones are formed by fixing the beam direction, responding only to audio in the beam direction to reduce the influence of interference factors. However, this still cannot solve the problem of large differences in audio volume caused by the varying volume of each participant's speech and their different distances from the audio acquisition component.
[0083] Therefore, this application proposes an audio processing method, electronic device, and storage medium to address the above-mentioned problems, which can solve the problem that the audio volume of the acquired audio data varies greatly due to the different speaking volumes of each participant and the different distances from the audio acquisition component.
[0084] Figure 1 This is a schematic diagram of an audio processing method system provided in one embodiment of this application. According to... Figure 1 As can be seen, the system 10 includes at least: an audio acquisition component 110, an image acquisition component 120, and an electronic device 130. This embodiment uses the electronic device 130 as a remote conferencing terminal (such as a video conferencing terminal or an audio conferencing terminal) as an example for illustration. In other embodiments, the electronic device 10 can also be a mobile phone, a laptop, a television, etc. This embodiment does not limit the implementation method of the electronic device.
[0085] The audio acquisition component 110 is used to acquire audio signals within a meeting setting (hereinafter referred to as the meeting room). The audio acquisition component 110 includes, but is not limited to, a microphone array or a sound intensity probe. Different audio acquisition components 110 are suitable for the same or different scenarios. This embodiment does not limit the implementation method of the audio acquisition component 110.
[0086] In this embodiment, the audio acquisition component 110 has a beamforming function. Beamforming is the process of forming a beam in a specified direction to acquire audio signals falling within the beam and suppress audio signals outside the beam.
[0087] For example, the audio acquisition component 110 forms a target beam towards the target sound source, causing the audio signal from the target sound source to fall within the target beam, thereby acquiring the audio signal from the target sound source and suppressing audio signals outside the target beam. Here, the target sound source refers to the sound source from which the user wants to acquire audio data, and the target beam is the beam pointing towards the target sound source.
[0088] Taking the audio acquisition component 110's working scenario as a remote conferencing scenario as an example, the target sound source includes both the participants attending the meeting and those who place their communication devices in the meeting room and participate in the meeting through those devices. This embodiment does not limit the implementation method of the target sound source.
[0089] The audio acquisition component 110 is communicatively connected to the electronic device 130 and sends the acquired audio signal to the electronic device 130.
[0090] Optionally, the audio acquisition component 110 may be a device independent of the electronic device 130, or it may be integrated into the electronic device 130. This embodiment does not limit the implementation of the audio acquisition component 110 and the electronic device 130.
[0091] Electronic device 130 is used to acquire audio signals and generate an acoustic image based on acoustic imaging principles. This acoustic image is used to display the audio information of each sound source within the venue. The audio information includes, but is not limited to, the relative positions of each sound source to the audio acquisition component 110, and the volume of each sound source.
[0092] The acoustic image includes the audio peaks of each sound source, which are visual representations of the maximum volume of the sound wave in the acoustic image. The position of the audio peak indicates the relative positional relationship between the sound source corresponding to the audio peak and the audio acquisition component 110.
[0093] The acoustic image is generated based on sound source localization technology using time-of-arrival (TOA). Specifically, the time difference between the arrival times of the audio signals from the sound sources at each pair of audio acquisition components 110 is measured to obtain a set of equations representing the sound source's position coordinates. Solving these equations yields the relative positional relationship between the sound source and the audio acquisition components 110. The amplitude of the sound source is also measured, and the spatial distribution of the sound sources is then visually represented to obtain the acoustic image. In the acoustic image, audio peaks represent each sound source; therefore, the position of the audio peaks indicates the relative positional relationship between each sound source and the audio acquisition components. In other embodiments, the acoustic image can also be generated based on beamforming technology, effectively forming a beam in the desired direction and picking up only the audio signal within the beam. This suppresses out-of-beam noise while extracting the audio signal within the beam. Correspondingly, the acoustic image can also be generated based on acoustic holography. This embodiment does not limit the generation method of the acoustic image.
[0094] Audio peaks can be represented using image markers with brightness and color distinctions, where different colors or brightness levels indicate different volumes; alternatively, audio peaks can be represented using contour lines, with sound sources of the same volume located on the same contour line and sound sources of different volumes located on different contour lines to indicate volume intensity. This embodiment does not limit the method of representing audio peaks.
[0095] Because there may be unwanted noise in the venue, such as the sound of air conditioning or projectors, these noises may also have corresponding audio peaks in the audiogram, i.e., spurious audio peaks as mentioned below. Therefore, it is necessary to remove spurious audio peaks from the audiogram to improve its signal-to-noise ratio (SNR). The SNR is the ratio of normal sound signal to noise; a higher SNR indicates less noise, and thus a higher quality audiogram.
[0096] Based on this, the audio processing method system provided in this embodiment also includes an image acquisition component 120. The image acquisition component 120 can acquire images based on the principle of optical imaging. In this way, the electronic device can combine the image to filter the audio peaks in the audio image, so as to improve the accuracy of the sound source location obtained through the audio image.
[0097] The image acquisition component 120 has a acquisition range that includes all sound sources corresponding to the audio peaks in the audio image. In one example, if the working scenario of the image acquisition component 120 is a remote conference scenario, since the various sound sources reflected in the audio image are usually located in the central area of the conference scenario, the image acquisition component 120 is installed at the edge of the conference scenario directly opposite the central area. In this case, the image acquired by the image acquisition component 120 is a panoramic image covering all sound sources in the entire conference scenario.
[0098] In this embodiment, the image acquisition component 120 is used to acquire images within the venue, enabling the electronic device 10 to obtain the first location information of the target sound source. The image acquisition component 120 includes, but is not limited to, panoramic cameras, scanners, etc. This embodiment does not limit the implementation method of the image acquisition component 120.
[0099] The first location information includes the target orientation and target distance of the target sound source relative to the image acquisition component 120.
[0100] The image acquisition component 120 is communicatively connected to the electronic device 130 and sends the acquired panoramic image to the electronic device 130.
[0101] Optionally, the image acquisition component 120 may be a device independent of the electronic device 130, or it may be integrated into the electronic device 130. This embodiment does not limit the implementation of the image acquisition component 120 and the electronic device 130.
[0102] Since the location and size of each target sound source in the venue are different, there may be distortion during the image acquisition process of the image acquisition component 120. Therefore, it may be inaccurate to determine the distance of the target sound source relative to the image acquisition component 120 by using only the image acquired by the image acquisition component 120.
[0103] To address the aforementioned issues, in this embodiment, the electronic device is used to: perform image recognition on images acquired by the image acquisition component within the meeting room to obtain first location information of the target sound source; generate an acoustic image based on the audio signal acquired by the audio acquisition component, the acoustic image being used to indicate the relative positional relationships between each sound source and the audio acquisition component within the meeting scene; determine the target relative positional relationship between the target sound source and the audio acquisition component using the first location information in each relative positional relationship; determine the target beam and volume gain of the target sound source based on the target relative positional relationship; and adjust the audio volume of the audio data acquired according to the target beam using the volume gain.
[0104] Generally, the farther the target sound source is from the audio acquisition component 110, the lower the volume of the acquired audio data. Therefore, the volume gain is positively correlated with the distance indicated by the relative position of the target. That is, the farther the target sound source is from the audio acquisition component 110, the greater the volume gain.
[0105] In this embodiment, the first position information obtained by the image acquisition component is used to filter out the relative position relationships corresponding to the target sound source from various relative position relationships. Therefore, even if there are interference factors such as noise, echo, and reverberation in the venue, the audio acquisition component can obtain the accurate position information of the target sound source and set the target beam and volume gain for the target sound source, so that the volume of the target sound sources will not differ too much. This can not only improve the quality of audio data in the meeting, but also make the audio volume of each target sound source consistent.
[0106] Furthermore, by combining the first position information and various relative positional relationships to determine the target relative positional relationship between the target sound source and the audio acquisition component, the problem of the target beam deviating from the target sound source due to inaccurate first position information caused by image distortion can be avoided. It can also avoid the problem of the target beam pointing to noise because noise can also form audio peaks in the sound image. By combining the first position information and various relative positional relationships to determine the target relative positional relationship corresponding to the target sound source, the target beam pointing to the target sound source can be accurately set. The audio volume of the audio data of the target sound source acquired in the target beam can be adjusted by adjusting the volume gain, so that the volume of each target sound source tends to be consistent and conforms to the user's listening habits, thereby improving the quality of the conference audio.
[0107] The audio processing method provided in this application will now be described in detail. The following embodiments use this method for... Figure 1The method described is specifically used in the processor of the electronic device. In actual implementation, the method can also be used in other devices that communicate with the electronic device, such as user terminals or servers. User terminals include, but are not limited to, mobile phones, computers, tablets, wearable devices, etc. This embodiment does not limit the implementation method of other devices or user terminals.
[0108] The communication connection can be wired or wireless. The wireless communication method can be short-range communication or wireless communication, etc. This embodiment does not limit the communication method between the self-moving device and other devices.
[0109] Figure 2 This is a flowchart of an audio processing method provided in one embodiment of this application. The method includes at least the following steps:
[0110] Step 201: Perform image recognition on the images acquired by the image acquisition component within the conference scene to obtain the first location information of the target sound source.
[0111] Image recognition refers to the technology of processing, analyzing, and understanding images acquired by image acquisition components.
[0112] The first location information includes: the target orientation and target distance between the image acquisition component and the target sound source; wherein the target orientation includes the horizontal angular orientation and the vertical angular orientation. In other embodiments, the target orientation may also include the vertical angular orientation and / or the horizontal angular orientation, and this embodiment does not limit the implementation method of the first location information.
[0113] Image recognition is performed on images captured by the image acquisition component within the conference scene to obtain the first location information of the target sound source, including: performing face detection on the image to obtain the face pixel height and face center point coordinates; obtaining the screen pixel size, screen center point coordinates, horizontal field of view, vertical field of view, and true face height of the image acquisition component, where the screen pixel size includes both horizontal and vertical dimensions; obtaining the horizontal and vertical distances between the face center point coordinates and the screen center point coordinates; using the horizontal dimension, horizontal distance, and horizontal field of view, obtaining the horizontal angular orientation between the face center point coordinates and the image acquisition component; using the vertical dimension, vertical distance, and vertical field of view, obtaining the vertical angular orientation between the face center point coordinates and the image acquisition component; using the vertical dimension and vertical field of view, obtaining the focal length; and using the focal length, true face height, and face pixel height, obtaining the target distance.
[0114] The actual height of the face is represented by the average height of the face in reality. For example, the average height of the face can be 25 cm, 23 cm, etc. This embodiment does not limit the actual height of the face.
[0115] Face detection is performed on images, including but not limited to at least one of the following face detection algorithms: Eigenfaces algorithm, Fisherfaces algorithm, etc.
[0116] Based on the face detection results of the image using a face detection algorithm, the target sound source in the image is determined.
[0117] Specifically, taking the image acquisition component at the center as an example, the target image acquired by the image acquisition component is as follows: Figure 3 As shown, using a neural network model, the target sound source 30 can be identified from sources such as air conditioner 32, door 31, projector 34, and target sound source 30.
[0118] To illustrate, in some special meetings, the types of attendees who can speak in the room are restricted. In such cases, the identified target sound sources usually need to be screened again. Therefore, in the process of identifying target sound sources in an image using a preset image recognition algorithm, it is also possible to identify the target type of the target sound source. The target type can be set by the user, and for example, the target type includes, but is not limited to: men, people aged 60 or older, standing people, or specific people, etc.
[0119] Traditional methods for obtaining target azimuth require calculation of the focal length, which is variable. However, the focal length calculated using a model may contain errors, affecting the accuracy of the target azimuth. Therefore, this embodiment provides a method for calculating target azimuth that does not require the use of focal length, as follows:
[0120] The horizontal angular orientation between the face center point coordinates and the image acquisition component is obtained using the horizontal dimension, horizontal distance, and horizontal field of view. Figure 4 As shown, the image acquisition component images the actual face 41 onto the imaging screen 43 through the lens 42, resulting in the following formula:
[0121]
[0122]
[0123] Based on the principle of triangle similarity, half of the horizontal field of view is the same as the azimuth angle β, and the horizontal azimuth angle is the same as the horizontal angle α. Therefore, α represents the horizontal azimuth angle, and β represents half of the horizontal field of view of the image acquisition component; fw is the horizontal distance between the coordinates of the center point of the face and the coordinates of the center point of the imaging screen 43; F is the focal length of the image acquisition component; and W is the horizontal dimension.
[0124] Therefore, using the ratio of formula (1) and formula (2), we can obtain the following formula:
[0125]
[0126] Since the focal length F is inaccurate, the horizontal azimuth α calculated by formula (1) alone is also inaccurate. However, the horizontal field of view, horizontal dimension, and horizontal distance are accurate. Therefore, formula (3) can be derived by combining formula (1) and formula (2), which can avoid using an inaccurate focal length in the calculation process. The horizontal angle α calculated by formula (3) can accurately characterize the horizontal azimuth between the target sound source and the image acquisition component.
[0127] Accordingly, by replacing the horizontal field of view, horizontal dimension, and horizontal distance with the vertical field of view, vertical dimension, and vertical distance, the accurate vertical angular orientation can be calculated based on the principle of triangle similarity and the above formula. This embodiment will not elaborate further here.
[0128] The target distance is obtained using focal length, actual face height, and face pixel height, such as... Figure 4 As shown, the following formula can be obtained:
[0129]
[0130] Among them, based on the principle of triangle similarity, half of the vertical field of view is the same as the azimuth angle γ, so γ represents half of the vertical field of view; H is the vertical distance.
[0131] Transforming formula (1), we obtain the following formula for calculating focal length F:
[0132]
[0133] like Figure 5 As shown, the image acquisition component images the actual face 41 onto the imaging screen 43 through the lens 42. Based on the principle of triangle similarity, the ratio of the actual height of the face to the target distance can be obtained as equal to the ratio of the face pixel height to the focal length, as shown in the following formula:
[0134]
[0135] Where h is the actual height of the face; d is the target distance; and hp is the face pixel height.
[0136] Combining formulas (5) and (6), we obtain the following formula:
[0137]
[0138] Based on formula (7), the target distance can be calculated.
[0139] Step 202: Generate an acoustic image based on the audio signal acquired by the audio acquisition component.
[0140] Step 203: Use the first position information to determine the target relative position relationship between the target sound source and the audio acquisition component in each relative position relationship.
[0141] The audio image contains the audio peaks corresponding to each sound source. These sound sources include the user's desired target sound source (such as attendees in a meeting room) and may also include unwanted noise. Therefore, it is necessary to extract the target peak corresponding to the target sound source from the audio peaks.
[0142] Because the image coordinate system used by the image acquisition component and the sound image coordinate system used by the audio acquisition component are different, the first position information needs to be mapped onto the sound image map in order to determine the target relative position relationship between the target sound source and the audio acquisition component in each relative position relationship.
[0143] In this embodiment, the target relative positional relationship between the target sound source and the audio acquisition component is determined using the first positional information in various relative positional relationships, including: obtaining the second positional information corresponding to the first positional information in the sound image; identifying audio peaks in each audio peak of the sound image that do not match the second positional information as false audio peaks; deleting false audio peaks from the audio peaks of the sound image to obtain the target relative positional relationship indicated by each deleted audio peak.
[0144] The second location information includes the first orientation and first distance between the audio acquisition component and the target sound source. The first orientation can be a horizontal angular orientation or a vertical angular orientation; this embodiment does not limit the implementation method of the first orientation.
[0145] The relative positional relationship of the target includes the second orientation and the second distance between the audio acquisition component and the target sound source. The second orientation is of the same type as the first orientation. For example, if the first orientation is a horizontal angular orientation, then the second orientation is also a horizontal angular orientation; if the first orientation is a vertical angular orientation, then the second orientation is also a vertical angular orientation.
[0146] In one example, the relative positions of the audio acquisition component and the image acquisition component in the venue are fixed. In this case, the coordinate transformation relationship between the image coordinate system and the audio coordinate system is fixed. Obtaining the second position information corresponding to the first position information in the audio-visual diagram includes: using the first position information to determine the relative position of the target sound source relative to the audio acquisition component in the image coordinate system; and using the coordinate transformation relationship between the image coordinate system and the audio coordinate system to transform this relative position to the audio coordinate system, thus obtaining the second position information.
[0147] Optionally, the line connecting the installation positions of the image acquisition component and the audio acquisition component is perpendicular to the horizontal plane. In this case, the relative position of the target sound source with respect to the image coordinate system is approximately the same as the relative position of the target sound source with respect to the audio acquisition component. Therefore, when using the first position information to determine the relative position of the target sound source with respect to the audio acquisition component in the image coordinate system, the first position information can be directly used as the relative position.
[0148] Alternatively, the installation position of the image acquisition component relative to the audio acquisition component is fixed. When using the first position information to determine the relative position of the target sound source relative to the audio acquisition component in the image coordinate system, the installation position relationship of the image acquisition component relative to the audio acquisition component is obtained, and the first position information is converted into the relative position of the target sound source relative to the audio acquisition component using the installation position relationship.
[0149] In other embodiments, the acquisition range of the image acquisition component may include the location of the audio acquisition component. In this case, the electronic device can also determine the relative position of the audio acquisition component with respect to the image acquisition component based on the position of the audio acquisition component in the image; this embodiment does not limit the method of determining this relative position.
[0150] Schematic, the image coordinate system is established based on the intermediate baseline of the image. For example, the midpoint of the intermediate baseline is used as the origin of the image coordinate system, the intermediate baseline is used as the y-axis of the image coordinate system, and the direction perpendicular to the intermediate baseline is used as the x-axis of the image coordinate system.
[0151] For example, with Figure 3 The intermediate baseline 35 of the image is used as a mapping reference. An image coordinate system is established using the intermediate baseline. In the image coordinate system, the first position information of each participant 30 relative to the image acquisition component can be determined.
[0152] Because distortion occurs during image acquisition, and because the true location and features of each target sound source are different, but the target orientation calculated by the above formula (3) is accurate, there is a slight difference between the target distance in the acquired image and the true distance information of the target sound source. Therefore, when converting the first position information to the second position information in the acoustic image, the first distance obtained by converting the target orientation may differ from the true distance information. Based on this, in this embodiment, the relative positional relationship of each sound source in the acoustic image can also be used to correct the first distance in the second position information.
[0153] In this embodiment, identifying false audio peaks whose relative positional relationship does not match the second positional information among the various audio peaks in the audio image includes at least the following steps:
[0154] Step 1: In the audio-visual image, obtain the distance difference between the second distance and the first distance.
[0155] Step 2: Identify false audio peaks where the second location does not match the first location and / or the distance difference is greater than a preset distance threshold.
[0156] The preset distance threshold can be pre-stored in the electronic device or input by the user. This embodiment does not limit the implementation method of the preset distance threshold.
[0157] For example, such as Figure 6 As shown, the second azimuth corresponding to the false audio peak 61 does not match the azimuth angles r1, r2, r3, 0°, 11°, 12°, and 13 corresponding to each target sound source, and the distance difference between the second distance corresponding to the false peak 61 and the first distance corresponding to each target sound source is also greater than the preset distance threshold. The false audio peak 61 is replaced by the false audio peak 62, and the same principle can be used to determine the false audio peak 62. In this embodiment, the azimuth angle is calculated with the vertical angle as 0°, moving left and right respectively. In other embodiments, it can also be calculated clockwise with the left horizontal line as 0°, or counterclockwise with the right horizontal line as 0°. This embodiment does not limit the implementation method of the azimuth angle.
[0158] Although the second azimuth corresponding to the false peak 63 matches the azimuth angle 0, the difference between the second distance corresponding to the false audio peak 63 and the distance to each target sound source is greater than the preset distance threshold.
[0159] Optionally, if the audio acquisition component is damaged or not all target sound sources have entered the venue, the number of audio peaks in the audio image acquired within the venue will be less than the number of target sound sources. Therefore, before identifying false audio peaks whose relative positional relationship does not match the second positional information among the various audio peaks in the audio image, it can be determined whether the number of audio peaks in the audio image is greater than or equal to the number of target sound sources.
[0160] In this embodiment, determining false audio peaks whose relative positional relationship does not match the second positional information among the various audio peaks in the audio image includes: when the number of audio peaks in the audio image is greater than or equal to the number of target sound sources, determining false audio peaks whose relative positional relationship does not match the second positional information among the various audio peaks in the audio image.
[0161] If the number of audio peaks in the audiogram is less than the number of target sound sources, the audiogram is re-acquired after a preset time interval until the number of audio peaks in the audiogram is greater than or equal to the number of target sound sources, or the number of times the audiogram is acquired reaches a preset threshold. The preset time interval and preset threshold can be pre-stored in the electronic device, input by the user, or obtained from other devices. This embodiment does not limit the method of obtaining the preset time interval and preset threshold.
[0162] Specifically, the preset time interval is 15 seconds and the preset number of times threshold is 3 times; or the preset time interval is 40 seconds and the preset number of times threshold is 2 times. Of course, the preset time interval and the preset number of times threshold can also be other values, and this embodiment does not limit the preset time interval and the preset number of times threshold.
[0163] If the number of times the audio-visual image is reacquired reaches a preset threshold, the electronic device is controlled to output an abnormal prompt for the audio component, such as text information or sound information. This embodiment does not limit the implementation method of the abnormal prompt.
[0164] In some special meetings, the types of attendees who can speak within the meeting room are restricted. In such cases, the audio peaks corresponding to target sound sources other than the target type should also be identified as false audio peaks. Therefore, the target relative positional relationship between the target sound source and the audio acquisition component is determined using the first positional information in various relative positional relationships. This includes: identifying false audio peaks in the various audio peaks of the albumen graph whose relative positional relationship does not match the first positional information of the target sound source of the target type; deleting the false audio peaks from the audio peaks of the albumen graph; and obtaining the target relative positional relationship indicated by the deleted audio peaks.
[0165] The method for determining false audio peaks where the relative positional relationship does not match the first positional information of the target sound source of the target type is described above, except that the target sound source is replaced with the target sound source of the target type. This embodiment will not be repeated here.
[0166] In other embodiments, determining the target relative positional relationship between the target sound source and the audio acquisition component using the first positional information in various relative positional relationships includes: determining, among the various audio peaks in the albumen graph, the actual audio peak that matches the second positional information of the target sound source of the target type, to obtain the target relative positional relationship indicated by the actual audio peak.
[0167] Among them, the true audio peak value refers to the audio peak value formed in the audio image of the target sound source that the user expects to obtain.
[0168] Correspondingly, distortion will occur during the image acquisition process, and because the true location and true features of each target sound source are different, the first location information in the acquired image will have slight differences from the true location information of the target sound source. Therefore, the second location information converted into the sonograph will also differ from the true location information.
[0169] In this embodiment, determining the true audio peak whose relative positional relationship matches the second positional information among the various audio peaks in the acoustic image includes at least the following steps:
[0170] Step 1: In the sonogram, the distance difference between the second distance and the first distance.
[0171] Step 2: Determine the true audio peak value that matches the first location in the second location and whose distance difference is less than or equal to a preset distance threshold.
[0172] Step 204: Determine the target beam and volume gain of the target sound source based on the relative position of the target.
[0173] Using the second orientation and second distance in the relative position relationship of the target to set the target beam corresponding to the target sound source, the audio acquisition component only acquires the sound within the target beam and suppresses the sound outside the target beam, which can reduce the interference of noise in the venue.
[0174] Specifically, such as Figure 7 As shown, different target beams 71 are set for target sound sources 30 with different orientations and distances. Different target beams 71 point to different target sound sources 30. The target beams 71 correspond one-to-one with the target sound sources 30. At this time, the audio acquisition component 110 only acquires the audio signal of the target sound source 30 in the direction of the target beam 71.
[0175] Volume gain is used to adjust the volume of audio data from each target sound source, making the volume of audio data from each target sound source more consistent, thus better conforming to the user's listening habits.
[0176] Schematic, determining the volume gain of a target sound source based on the relative positional relationship of the targets includes: determining an intermediate distance using the maximum and minimum distances indicated by the relative positional relationship of the targets; obtaining a standard volume gain corresponding to the intermediate distance; determining the gain difference between the volume gain of each target sound source and the standard volume gain based on the difference between each target sound source and the intermediate distance; and determining the volume gain of each target sound source based on the standard volume gain and the gain difference.
[0177] The maximum and minimum distances are obtained by pairwise comparison among all distances indicated by relative positional relationships. This embodiment does not limit the method of obtaining the maximum and minimum distances.
[0178] Determining the median distance using the maximum and minimum distances involves finding the median between the maximum and minimum distances to obtain the median distance.
[0179] The standard volume gain corresponding to the intermediate distance is pre-stored in the electronic device. This standard volume gain can be input by the user or obtained from other devices. This embodiment does not limit the method of obtaining the standard volume gain.
[0180] The standard volume gain can be set to 12 dB or 6 dB. Of course, the standard volume gain can be other values. This embodiment does not limit the value of the preset standard volume.
[0181] The difference between the distance to other target sound sources and the intermediate distance is obtained by relative positional relationship, and the volume gain of other target sound sources is allocated according to the difference and standard volume gain.
[0182] In one example, the difference between the distance to other target sound sources and the median distance is obtained through relative positional relationships. The volume gain of other target sound sources is then allocated based on this difference and the standard volume gain, expressed by the following formula:
[0183]
[0184] Among them, D mid Indicates the intermediate distance, Gbase_mid represents the standard volume gain, and n is the code of the target beam. The coding order of the target beam is as follows: Figure 8 As shown, the target beams are sequentially encoded from left to right as 1, 2, ..., 7. In other embodiments, the target beams may have other encoding orders. This embodiment does not limit the encoding order of the target beams; Gbase_n refers to the volume gain of the nth target beam; D n This is the second distance between the target sound source corresponding to the nth target beam and the audio acquisition component.
[0185] In another example, the difference between the distance to other target sound sources and the median distance is obtained through relative positional relationships. The volume gain of other target sound sources is then allocated based on this difference and the standard volume gain. This further includes obtaining the difference between the distance to other target sound sources and the median distance. This difference can be positive or negative. When the difference is positive, it indicates that the distance to the target sound source is greater than the median distance, and a volume gain greater than the standard volume gain should be allocated. When the difference is positive, it indicates that the distance to the target sound source is less than the median distance, and a volume gain less than the standard volume gain should be allocated. In this case, the standard volume gain and the difference are positively correlated. In actual implementation, the volume gain of other target sound sources can also be allocated based on other differences and the standard volume gain; this embodiment does not limit this approach.
[0186] Different differences correspond to different volume gains, and the relationship between the differences and volume gains is pre-stored in the electronic device.
[0187] Because the volume of the target sound source changes constantly when it is emitting sound, even if the volume gain is set to make the volume of the target sound source more consistent, there is still a problem that the adjusted volume is too high or too low, making the speech of the target sound source unclear.
[0188] Optionally, determining the volume gain of the target sound source based on the relative positional relationship of the target includes: determining a first gain value of the target sound source based on the relative positional relationship of the target; determining the volume difference between the audio volume obtained by adjusting the audio data according to the first gain value and the preset standard volume; using the volume difference to determine a second gain value of the target sound source; and determining the volume gain of the target sound source based on the first gain value and the second gain value.
[0189] Among them, the first gain value is positively correlated with the distance indicated by the relative position of the target.
[0190] Determining the volume difference between the audio volume obtained by adjusting the audio data according to the first gain value and the preset standard volume includes: determining the difference between the audio volume of the target sound source audio data after adjustment according to the first gain and the preset standard volume, and obtaining the volume difference.
[0191] The preset standard volume indicates a volume that allows users to clearly hear the target sound source without being overly harsh, conforming to the user's listening habits. The preset standard volume can vary depending on the meeting. For example, in a meeting where discussion is permitted, the volume of participants' discussions will affect the volume they can hear; in this case, the preset standard volume could be 60 decibels or another value. This embodiment does not limit the value of the preset standard volume. The preset standard volume can be set by the user or it can be a default value stored in the electronic device; this embodiment does not limit the method of setting the preset standard volume.
[0192] In actual implementation, volume difference can also be obtained in other ways, such as the ratio between audio volume and standard volume. This embodiment does not limit the implementation method of volume difference.
[0193] Using volume differences to determine the second gain value of the target sound source allows for the addition of a second gain value to the first gain value, ensuring the adjusted volume is neither too loud nor too soft, thus better aligning with the user's listening habits. Optionally, since the target sound source in the venue is not stationary, the relative position between the target sound source and the audio acquisition component may change. In this case, the target beam may fail to acquire the audio signal from the target sound source. Therefore, after a preset time interval, the electronic device is controlled to re-execute steps 201 to 204 to ensure the target beam changes in time to follow the movement of the target sound source, enabling the audio acquisition component to accurately and promptly set and adjust the target beam.
[0194] Step 205: Use volume gain adjustment to adjust the audio volume of the audio data acquired according to the target beam.
[0195] Since there are speakers and ordinary attendees in the meeting room, the speaker is usually responsible for explaining the important content of the meeting or driving the meeting process. Therefore, when the speaker speaks, the speaker's voice should be emphasized as much as possible to avoid the speaker's voice being drowned out by the discussion of the attendees, which would delay the progress of the meeting.
[0196] Alternatively, if there are attendees who interact with the speaker, their remarks are usually related to the speaker's remarks, that is, related to the important content of the meeting. Therefore, their voices should also be highlighted as much as possible.
[0197] Alternatively, during a meeting, participants may have private discussions about certain issues. In this case, the sound of these private discussions can be considered noise. The sound is usually low, but after being amplified, it can affect other participants who are speaking.
[0198] Therefore, it is necessary to determine whether the target sound source meets the preset volume adjustment conditions to determine whether the volume of the participants should be adjusted. If the target sound source meets the volume adjustment conditions, the step of adjusting the volume of the audio data acquired according to the target beam using volume gain is triggered; if the target sound source does not meet the volume adjustment conditions, and the volume of the audio data is greater than or equal to the preset volume threshold, the audio volume of the audio data is suppressed.
[0199] The volume adjustment conditions include at least one of the following: the target sound source is speaking; the target sound source is the speaker; or the target sound source is a participant interacting with the speaker.
[0200] Specifically, methods for determining whether a target sound source is in a speaking state include, but are not limited to: determining whether the target sound source is in a speaking state based on the opening and closing state of the target sound source's lips; or, determining whether the target sound source is in a speaking state based on the facial area of the target sound source; or, determining whether the target sound source is the speaker based on the identity information of the target sound source; or, determining whether the participant is interacting with the speaker based on posture information. In actual implementation, other methods can also be used to determine whether the target sound source meets the preset volume adjustment conditions. This embodiment does not limit the method of determining the volume adjustment conditions.
[0201] The identity information can be entered by the user or obtained from other devices. Identity information refers to information that can indicate whether a participant is the main speaker. Identity information includes, but is not limited to, gender, name, special markings, etc.
[0202] Optionally, suppressing audio volume may include, but is not limited to, reducing the volume gain to 0; or, using other methods to directly reduce the audio volume of the audio data corresponding to the target sound source.
[0203] The audio processing method provided in this embodiment will be described below with a specific example. In this example, the target types are participants and speakers. Figure 9 As shown.
[0204] Step 901, Acoustic imaging and optical imaging coordinate mapping, converting the first position information determined based on the image acquired by the image acquisition component into the second position information in the acoustic image.
[0205] Step 902: Combine the second position information to delete false audio peaks in the sound image to obtain the target relative position relationship between the target sound source and the audio acquisition component; use the target relative position relationship to determine the intermediate distance between the target sound source and the audio acquisition component, and set the intermediate distance as the standard volume gain of the first gain value; use the difference between the distance between each target sound source and the audio acquisition component and the intermediate distance to determine the gain difference and obtain the volume gain of each target sound source.
[0206] Step 903: Determine whether the meeting has ended. If the meeting has ended, the process ends. If the meeting has not ended, determine whether the target sound source is in a speaking state. If the target sound source is in a speaking state, proceed to step 904. If the target sound source is not in a speaking state, proceed to step 915.
[0207] Step 904: Determine the target type of the target sound source; if the target type of the target sound source is determined to be the speaker, proceed to step 905; if the target type of the target sound source is determined to be the attendee, proceed to step 907.
[0208] Step 905: Compare the audio volume of the speaker's audio data with the preset standard volume to obtain the volume difference; set a second gain value based on the volume difference; add the beam information of the target beam corresponding to the speaker to the beam protection pool.
[0209] The beam protection cell stores beam information of the target beam for which volume adjustment is required. This beam information can be the target beam's number or the target sound source's account number, etc. This embodiment does not limit the implementation method of the beam information.
[0210] Step 906: Reduce the audio volume of the unprotected audio data collected in the target beam outside the beam protection pool to highlight the audio volume of the protected audio data collected in the target beam inside the beam protection pool; then proceed to step 903.
[0211] Optionally, there must be at least a 6-decibel difference between the audio volume of unprotected audio data and the audio volume of protected audio data. The first-order volume during audio playback is 3 decibels, and adjustments to the first-order volume are not easily noticeable in a meeting room and will still be drowned out by the participants' speeches. Therefore, there must be at least two volume levels, i.e., a 6-decibel difference.
[0212] Step 907: Add the beam information of the target beam corresponding to the participant to the participant speech statistics pool; proceed to step 908.
[0213] The participant speaking statistics pool is used to store the beam information of the target beam corresponding to the participant who is speaking. Since steps 905 and 907 are mutually exclusive for the same target sound source, the beam information of the target beam corresponding to the speaker will not be added to the participant speaking statistics pool.
[0214] Step 908: Compare the audio volume of the target sound source in the participant speaking statistics pool with the preset standard volume to obtain the volume difference, and set the second gain value based on the volume difference; then proceed to step 909.
[0215] Step 909: Determine whether any participant in the participant speaking statistics pool has interacted with the speaker; if the participant in the participant speaking statistics pool has interacted with the speaker, proceed to step 910; if the participant in the participant speaking statistics pool has not interacted with the speaker, proceed to step 911.
[0216] Step 910: Add the beam information of the target beam corresponding to the participant to the beam protection pool; proceed to step 906.
[0217] The target beams corresponding to the participants in the beam protection pool are all moved in from the participant speech statistics pool.
[0218] Step 911: Determine whether the target beam corresponding to the participant is in the beam protection pool. If it is, remove the target beam from the beam protection pool; proceed to step 912.
[0219] Step 912: Determine whether the participants outside the beam protection pool are in a speaking state; if the participants outside the pool are in a speaking state, proceed to step 913; if the participants outside the pool are not in a speaking state, proceed to step 906.
[0220] Step 913: Determine whether the volume of the participants outside the beam protection pool has reached the preset volume threshold; if the volume of the participants outside the beam protection pool has reached the preset volume threshold, proceed to step 914; if the volume of the participants outside the beam protection pool has not reached the preset volume threshold, proceed to step 903.
[0221] Step 914: Attenuate the volume of participants outside the beam protection pool to below a preset volume threshold; proceed to step 903.
[0222] Step 915: Determine the target type of the target sound source; if the target sound source is determined to be the speaker, proceed to step 916; if the target sound source is determined to be a participant, proceed to step 917.
[0223] Step 916: Cancel the reduction of audio volume of unprotected audio data collected in the target beam outside the beam protection pool, and determine whether the beam information of the target beam corresponding to the speaker is in the beam protection pool. If the beam information of the target beam corresponding to the speaker is in the beam protection pool, then clear the beam information out of the beam protection pool and execute step 903.
[0224] Step 917: First, determine whether the beam information of the target beam corresponding to the participant is in the participant speaking statistics pool. If the beam information of the target beam corresponding to the participant is in the participant speaking statistics pool, then clear the beam information from the participant speaking statistics pool, then reset the target beam and volume gain corresponding to the participant, and execute step 903.
[0225] In summary, the audio processing method provided in this embodiment obtains the first location information of the target sound source by performing image recognition on the images captured by the image acquisition component in the conference scene; generates a sound image map based on the audio signals acquired by the audio acquisition component, which is used to indicate the relative positional relationships between each sound source and the audio acquisition component in the conference scene; determines the target relative positional relationship between the target sound source and the audio acquisition component using the first location information; determines the target beam and volume gain of the target sound source based on the target relative positional relationship; and adjusts the audio volume of the audio data acquired according to the target beam using the volume gain. This method can solve the problem of significant differences in sound between participants due to differences in the volume of each participant's voice and the distance between each participant and the sensor. Since the sound image map is used to indicate the relative positional relationships between each sound source and the audio acquisition component in the conference scene, this method can solve the problem of significant differences in sound between participants due to differences in the volume of each participant's voice and the distance between each participant and the sensor. The image shows the relative positions of various sound sources within the meeting room, including those affected by noise, echoes, reverberation, and other interference. Using an image can identify the target sound source and obtain its initial position information. Therefore, by using this initial position information to determine the target's relative position among various relative positions, even when noise, echoes, and reverberation are present in the audio-visual image, the relative positions of these interference factors can be eliminated, thus obtaining the target sound source's relative position. Since the relative position information obtained from the audio-visual image is relatively accurate, the precise position information of the target sound source can be obtained. Target beamforming and volume gain can then be set for the target sound source, ensuring that the volume differences between target sound sources are not too large. This improves the quality of audio data within the meeting and ensures consistent audio volume across all target sound sources.
[0226] Meanwhile, by combining the first position information and various relative positional relationships to determine the target relative positional relationship between the target sound source and the audio acquisition component, the problem of the target beam deviating from the target sound source due to inaccurate first position information caused by image distortion can be avoided; it can also avoid the problem of the target beam pointing to noise due to the formation of audio peaks in the sound image caused by noise; by combining the first position information and various relative positional relationships to determine the target relative positional relationship corresponding to the target sound source, the target beam pointing to the target sound source can be accurately set, and the audio volume of the audio data of the target sound source acquired in the target beam can be adjusted by adjusting the volume gain, so that the volume of each target sound source tends to be consistent and conforms to the user's listening habits, thereby improving the quality of conference audio.
[0227] At the same time, by using volume gain adjustment to adjust the audio gain of the audio data of the target sound source collected within the target beam, the problem of participants being unable to hear clearly due to excessively loud or soft volume can be avoided. Therefore, the number of times participants have to repeat themselves because they cannot hear clearly can be reduced, thus shortening meeting time and improving meeting efficiency.
[0228] In addition, by deleting false audio peaks that do not match the first position information from the audio image, the relative positional relationship of the target indicated by each audio peak after deletion can be obtained. This can improve the accuracy of the obtained relative positional relationship of the target, thereby improving the accuracy of the target beam pointing to the target sound source and improving the effectiveness of volume gain adjustment of the target sound source volume, thereby reducing the interference of noise and other interference factors in the meeting room on the meeting.
[0229] In addition, when the number of audio peaks is greater than or equal to the number of target sound sources, identifying false audio peaks whose relative positional relationship does not match the first positional information among the various audio peaks in the audio image can avoid errors in obtaining the target relative positional relationship of the target sound sources when the audio components are damaged or not all target sound sources have entered the conference scene, thereby improving the reliability of electronic devices.
[0230] In addition, by acquiring the first location information of the target sound source of the target type and identifying false audio peaks whose relative positional relationship does not match the first location information of the target sound source of the target type, different target sound sources can be flexibly determined according to different meeting scenarios, which can improve the flexibility of electronic devices.
[0231] In addition, by determining the volume gain of each target sound source based on the standard volume gain corresponding to the intermediate distance, the volume of the target sound sources can be made more consistent, which conforms to the user's listening habits. At the same time, it can also reduce the preset correspondence between distance and volume gain, saving storage resources of electronic devices.
[0232] In addition, by determining the second gain value based on the preset standard volume on the basis of the first gain value, the corresponding volume gain can be adjusted in a timely manner when the volume of the target sound source changes, and the volume change can be smoothly transitioned, thereby improving the accuracy of volume adjustment of electronic devices.
[0233] In addition, when the preset volume adjustment conditions are met, using volume gain adjustment to adjust the volume of the audio data collected according to the target beam can ensure that the volume of the target sound source that meets the volume adjustment conditions is highlighted, thereby improving the intelligence of volume adjustment of electronic devices.
[0234] In addition, suppressing the audio volume of the target sound source's audio data when the preset volume adjustment conditions are not met can prevent the target sound source from becoming too loud and interfering with the normal progress of the meeting, thus improving the efficiency of the meeting.
[0235] Figure 10 This is a block diagram of an apparatus for an audio processing method provided in one embodiment of this application. The apparatus includes at least the following modules: a first position module 1010, a sound image generation module 1020, a relative position module 1030, a volume gain module 1040, and a volume adjustment module 1050.
[0236] The first position module 1010 is used to perform image recognition on the images acquired by the image acquisition component in the conference scene to obtain the first position information of the target sound source.
[0237] The audio-visual generation module 1020 is used to generate an audio-visual image based on the audio signal acquired by the audio acquisition component.
[0238] The relative position module 1030 uses the first position information to determine the target relative position relationship between the target sound source and the audio acquisition component in various relative position relationships.
[0239] The volume gain module 1040 determines the target beam and volume gain of the target sound source based on the relative position of the target.
[0240] The volume adjustment module 1050 uses volume gain to adjust the audio volume of the audio data acquired according to the target beam.
[0241] For relevant details, please refer to the above method implementation examples.
[0242] It should be noted that the audio processing method apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when processing audio data. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the audio processing method apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing method apparatus provided in the above embodiments and the audio processing method embodiments belong to the same concept, and its specific implementation process can be found in the method embodiments, which will not be repeated here.
[0243] Figure 11 This is an electronic device provided in one embodiment of this application. The electronic device may be... Figure 1 The electronic device or other device that is communicatively connected to the electronic device includes at least a processor 1101 and a memory 1102.
[0244] Processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1101 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1101 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1101 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0245] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 are used to store at least one instruction, which is executed by the processor 1101 to implement the audio processing method provided in the method embodiments of this application.
[0246] In some embodiments, the electronic device may also optionally include a peripheral device interface and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to, radio frequency circuits, touch displays, audio circuits, and power supplies.
[0247] Of course, electronic devices may also include fewer or more components, and this embodiment does not limit this.
[0248] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the audio processing method of the above method embodiments.
[0249] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the audio processing method of the above-described method embodiments.
[0250] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0251] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An audio processing method, characterized in that, The method includes: Image recognition is performed on the images acquired by the image acquisition component within the conference scene to obtain the first location information of the target sound source; An acoustic image is generated based on the audio signal acquired by the audio acquisition component. The acoustic image is used to indicate the relative positional relationship between each sound source and the audio acquisition component in the conference scene. The acoustic image includes the audio peak value of each sound source, and the position of the audio peak value is used to indicate the relative positional relationship. Determining the target relative positional relationship between the target sound source and the audio acquisition component using the first positional information in various relative positional relationships includes: obtaining second positional information corresponding to the first positional information in the sound image; determining the real audio peaks that match the second positional information and the false audio peaks that do not match in each audio peak of the sound image, wherein the real audio peaks refer to the audio peaks formed by the target sound source in the sound image that the user expects to obtain; deleting the false audio peaks from the audio peaks in the sound image to obtain the target relative positional relationship indicated by each deleted audio peak; The target beam and volume gain of the target sound source are determined based on the relative positional relationship of the targets, wherein the volume gain is positively correlated with the distance indicated by the relative positional relationship of the targets; The volume gain is used to adjust the audio volume of the audio data acquired according to the target beam.
2. The method according to claim 1, characterized in that, The first location information includes the target orientation and the target distance, wherein the target orientation includes the horizontal angular orientation and the vertical angular orientation; The step of performing image recognition on images acquired by the image acquisition component within the conference scene to obtain the first location information of the target sound source includes: Face detection is performed on the image to obtain the face pixel height and the coordinates of the face center point; The screen pixel size, screen center point coordinates, horizontal field of view, vertical field of view, and true face height of the image acquisition component are obtained. The screen pixel size includes both horizontal and vertical dimensions. Obtain the horizontal and vertical distances between the coordinates of the face center point and the coordinates of the screen center point; The horizontal angular orientation between the face center point coordinates and the image acquisition component is obtained using the horizontal dimension, the horizontal distance, and the horizontal field of view. The vertical angular orientation between the face center point coordinates and the image acquisition component is obtained using the vertical dimension, the vertical distance, and the vertical field of view. The focal length is obtained using the vertical dimension and the vertical field of view. The target distance is obtained using the focal length, the actual height of the face, and the pixel height of the face.
3. The method according to claim 1, characterized in that, The second location information includes a first distance and a first orientation between the audio acquisition component and the target sound source, and the relative positional relationships include a second distance and a second orientation; The step of identifying false audio peaks whose relative positional relationship does not match the second positional information among the various audio peaks in the acoustic image includes: In the various audio peaks of the acoustic image, the distance difference between the second distance and the first distance is obtained. The false audio peak is determined to be one where the second orientation does not match the first orientation and / or the distance difference is greater than a preset distance threshold.
4. The method according to claim 1, characterized in that, The step of identifying false audio peaks whose relative positional relationship does not match the second positional information among the various audio peaks in the acoustic image includes: If the number of audio peaks in the sound image is greater than or equal to the number of target sound sources, false audio peaks whose relative positional relationship does not match the second positional information are identified among the audio peaks in the sound image.
5. The method according to claim 1, characterized in that, Determining the volume gain of the target sound source based on the relative positional relationship of the targets includes: The intermediate distance is determined using the maximum and minimum distances indicated by the relative positional relationship of the targets; Obtain the standard volume gain corresponding to the intermediate distance; Based on the difference between each target sound source and the intermediate distance, the gain difference between the volume gain of each target sound source and the standard volume gain is determined; The volume gain of each target sound source is determined based on the standard volume gain and the gain difference.
6. The method according to claim 1, characterized in that, Determining the volume gain of the target sound source based on the relative positional relationship of the targets includes: Based on the relative positional relationship of the targets, a first gain value of the target sound source is determined, wherein the first gain value is positively correlated with the distance indicated by the relative positional relationship of the targets; Determine the volume difference between the audio volume obtained by adjusting the audio data according to the first gain value and the preset standard volume; The volume difference is used to determine the second gain value of the target sound source; The volume gain of the target sound source is determined based on the first gain value and the second gain value.
7. The method according to claim 1, characterized in that, The method includes: Determine whether the target sound source meets preset volume adjustment conditions; the volume adjustment conditions include at least one of the following: the target sound source is in a speaking state; the target sound source is a speaker; the target sound source is a participant interacting with the speaker; If the target sound source meets the volume adjustment conditions, the step of adjusting the volume of the audio data acquired according to the target beam using the volume gain is triggered.
8. The method according to claim 7, characterized in that, The method further includes: If the target sound source does not meet the volume adjustment conditions, and the volume of the audio data is greater than or equal to a preset volume threshold, then the audio volume of the audio data is suppressed.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory connected to the processor, the memory storing a program, which the processor executes to implement the audio processing method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed by a processor, is used to implement the audio processing method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Video call equipment and audio gain method
CN112423191A
Optical image and acoustic image fusion method and system
CN112466323A