An audio judgment method and device, electronic equipment and storage medium

CN117956376BActive Publication Date: 2026-09-08GUANGZHOU KINDLINK SOFTWARE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211336717.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2026-09-08
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

[0003]目前,判断齐声朗读状态可以用基于深度学习的方法来实现,然而此类方法训练出来的模型和需要的计算量较大,在低性能的嵌入式设备上难以运行,提高了对齐声朗读的音频确认的成本

Benefits of technology

[0015] This application's embodiments acquire audio signals collected in a classroom; based on the audio signals, obtain sound arrival angle information and sound energy data of the audio signals, wherein the sound arrival angle information includes several sound arrival angles; if the sound energy data of the audio signals exceeds a preset sound energy threshold, obtain a preset interference sound arrival angle; based on the preset interference sound arrival angle, exclude the interference sound arrival angles in the sound arrival angle information to obtain excluded sound arrival angle information; if the number of sound arrival angles in the excluded sound arrival angle information exceeds a preset first sound source number threshold, determine that the audio signal is a unison reading, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to unison reading. By calculating the sound arrival angles and sound energy corresponding to the audio signals, accurate determination of unison reading audio is achieved, reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117956376B_ABST
    Figure CN117956376B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio judgment method and device, electronic equipment and storage medium, wherein the method comprises: acquiring an audio signal collected in a classroom; obtaining sound arrival angle information of the audio signal and sound energy data of the audio signal according to the audio signal; if the sound energy data of the audio signal exceeds a preset sound energy threshold, obtaining a preset interference sound arrival angle; excluding, according to the interference sound arrival angle, a sound arrival angle associated with the interference sound arrival angle from the sound arrival angle information, to obtain the sound arrival angle information after exclusion; and if the number of sound arrival angles in the sound arrival angle information after exclusion exceeds a preset first sound source number threshold, determining that the audio signal is a chorus reading. The embodiments of the present application realize accurate judgment of the audio of chorus reading and reduce the cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an audio judgment method, apparatus, electronic device and storage medium. Background Technology

[0002] In remote teaching scenarios, the audio captured by the microphone needs to be noise-reduced. Generally, deep learning-based noise reduction methods (such as RNNoise) are used. However, the audio characteristics of the audio signal in the state of reading aloud in unison are very similar to the characteristics of noise. The audio signal in the state of reading aloud in unison is often treated as noise and suppressed. Therefore, it is necessary to determine whether the current state is reading aloud in unison.

[0003] Currently, judging the state of unison reading can be achieved using deep learning-based methods. However, the models trained by such methods require a large amount of computation, making them difficult to run on low-performance embedded devices, which increases the cost of audio verification of unison reading. Summary of the Invention

[0004] Based on this, this application provides an audio judgment method, device, electronic device, and storage medium, which realizes accurate judgment of aligned sound reading by calculating the sound angle and sound energy corresponding to the audio signal, thereby reducing costs.

[0005] As a first aspect of the embodiments of this application, an audio determination method is provided, including the following steps:

[0006] Acquire audio signals collected in the classroom; based on the audio signals, obtain the sound arrival angle information and sound energy data of the audio signals, wherein the sound arrival angle information includes several sound arrival angles;

[0007] If the sound energy data of the audio signal exceeds a preset sound energy threshold, a preset interference sound arrival angle is obtained; based on the preset interference sound arrival angle, the interference sound arrival angle in the sound arrival angle information is excluded to obtain the excluded sound arrival angle information;

[0008] If the number of sound arrival angles after exclusion exceeds a preset first sound source number threshold, the audio signal is determined to be unison reading, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to unison reading.

[0009] As a second aspect of the embodiments of this application, an audio determination device is provided, including:

[0010] The data acquisition module is used to acquire audio signals collected in the classroom; based on the audio signals, it obtains the sound arrival angle information and the sound energy data of the audio signals, wherein the sound arrival angle information includes several sound arrival angles;

[0011] The sound arrival angle processing module is used to obtain a preset interference sound arrival angle if the sound energy data of the audio signal exceeds a preset sound energy threshold; and to exclude the interference sound arrival angle from the sound arrival angle information based on the preset interference sound arrival angle to obtain the excluded sound arrival angle information.

[0012] The target audio signal determination module is used to determine that the audio signal is a choral reading if the number of sound arrival angles of the excluded sound arrival angle information exceeds a preset first sound source number threshold, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to the choral reading.

[0013] As a third aspect of the present application, an electronic device is provided, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the audio judgment method as described in the first aspect.

[0014] As a fourth aspect of the present application, a storage medium is provided, the storage medium storing a computer program, which, when executed by a processor, implements the steps of the audio judgment method as described in the first aspect.

[0015] This application's embodiments acquire audio signals collected in a classroom; based on the audio signals, obtain sound arrival angle information and sound energy data of the audio signals, wherein the sound arrival angle information includes several sound arrival angles; if the sound energy data of the audio signals exceeds a preset sound energy threshold, obtain a preset interference sound arrival angle; based on the preset interference sound arrival angle, exclude the interference sound arrival angles in the sound arrival angle information to obtain excluded sound arrival angle information; if the number of sound arrival angles in the excluded sound arrival angle information exceeds a preset first sound source number threshold, determine that the audio signal is a unison reading, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to unison reading. By calculating the sound arrival angles and sound energy corresponding to the audio signals, accurate determination of unison reading audio is achieved, reducing costs.

[0016] To better understand and implement this application, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0017] Figure 1 A flowchart illustrating an audio determination method provided in one embodiment of this application;

[0018] Figure 2 This is a flowchart illustrating step S1 of an audio determination method provided in one embodiment of this application;

[0019] Figure 3 This is a flowchart illustrating step S2 of an audio determination method provided in one embodiment of this application;

[0020] Figure 4 This is a flowchart illustrating step S2 of an audio determination method provided in one embodiment of this application;

[0021] Figure 5 This is a schematic diagram of the structure of an audio determination device provided in one embodiment of this application;

[0022] Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Wherein, when the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0024] It should be understood that the embodiments described below do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0025] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, in the description of this application, unless otherwise stated, “a plurality” means two or more. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more associated listed items, for example, A and / or B, which can represent: A alone, A and B together, and B alone; the character “ / ” generally indicates that the preceding and following objects are in an “or” relationship.

[0026] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, this information should not be limited to these terms, and these terms are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances. Depending on the context, the word "if" as used in this application can be interpreted as "when," "when," or "in response to determination."

[0027] The application scenarios of the audio judgment method in this application include recording and broadcasting equipment, audio acquisition equipment, and loudspeakers; the audio acquisition equipment transmits data with the recording and broadcasting equipment, the audio acquisition equipment acquires the audio signal emitted by the loudspeaker and the audio signal emitted by the teacher or student, and then sends the acquired audio signal to the recording and broadcasting equipment, the recording and broadcasting equipment receives the audio signal, and performs noise reduction and saves the audio signal.

[0028] The audio judgment method of this application can be executed by a recording and broadcasting device. This recording and broadcasting device can implement the audio judgment method through software and / or hardware. The recording and broadcasting device can consist of two or more physical entities, or it can consist of a single physical entity. The hardware referred to as the recording and broadcasting device essentially refers to computer equipment; for example, the recording and broadcasting device can be a computer, mobile phone, tablet, or smart interactive whiteboard, or other smart devices.

[0029] The recording and broadcasting device is equipped with at least one type of operating system, including but not limited to Android, Linux, and Windows. In one embodiment, the recording and broadcasting device may install at least one application based on the operating system; in this embodiment, a meeting scheduling application is used as an example. This application may be a built-in application of the operating system or an application downloaded from a third-party device or server. Users can implement the audio judgment method of this application based on this meeting scheduling application.

[0030] Please see Figure 1 , Figure 1 The following is a flowchart illustrating an audio determination method provided in one embodiment of this application. The method includes the following steps:

[0031] S1: Acquire the audio signal collected in the classroom; based on the audio signal, obtain the sound arrival angle information and the sound energy data of the audio signal, wherein the sound arrival angle information includes several sound arrival angles.

[0032] The recording and broadcasting equipment can acquire audio signals sent by the audio acquisition equipment through data transmission, or it can extract corresponding audio signals from a preset database. This database includes audio signals collected at several different times during class. In an optional embodiment, the user can establish a connection between the recording and broadcasting equipment and a server by building a database, thereby enabling database data transmission between the recording and broadcasting equipment and the server. This allows the recording and broadcasting equipment to access the database and extract corresponding audio signals.

[0033] The recording and broadcasting equipment collects audio signals from the classroom through audio acquisition devices, obtains the audio signals collected in the classroom, and obtains the sound angle of arrival information of the audio signals based on the audio signals. The sound angle of arrival information includes several sound angles of arrival (DOA), and each sound angle of arrival corresponds to a sound source.

[0034] Recording and broadcasting equipment obtains sound energy data of audio signals. Specifically, the recording and broadcasting equipment can obtain the sound energy corresponding to the audio signal by detecting the amplitude generated by the audio acquisition device for the audio signal.

[0035] S2: If the sound energy data of the audio signal exceeds the preset sound energy threshold, obtain the preset angle of arrival of the interfering sound; based on the preset angle of arrival of the interfering sound, exclude the angle of arrival of the interfering sound in the angle of arrival information to obtain the angle of arrival information of the excluded sound.

[0036] Because the sound energy of the audio signal collected in the classroom during choral reading is much greater than that of normal speech, and there may be interfering sound sources in the audio signal, such as a fixed loudspeaker in the classroom, through which the teacher speaks, the sound energy of the audio signal emitted by that loudspeaker is also much greater than that of normal speech, the recording equipment has a preset sound energy threshold. The preset sound energy threshold is used to determine whether the sound energy of the audio signal collected by the audio acquisition device is the audio signal of choral reading or the audio signal emitted by the teacher through the loudspeaker. The recording equipment can make an initial judgment on the audio signal based on the sound energy of the audio signal collected by the audio acquisition device to determine whether the sound energy of the audio signal collected by the audio acquisition device is the audio signal of choral reading or the audio signal emitted by the teacher through the loudspeaker. Specifically, the preset sound energy threshold can be set according to the actual situation and is not limited thereto.

[0037] In this embodiment, the recording device compares the sound energy corresponding to the audio signal with a preset sound energy threshold. If the sound energy corresponding to the audio signal does not exceed the sound energy threshold, it is determined that the audio signal does not include the audio signal in the state of reading aloud in unison, nor does it include the audio signal emitted by the teacher through the speaker. Therefore, the audio signal can be directly processed for noise reduction.

[0038] If the sound energy data of the audio signal exceeds the preset sound energy threshold, the sound energy of the audio signal collected by the audio acquisition device may only include the audio signal in the state of reading aloud in unison, or it may only include the audio signal emitted by the teacher through the speaker, or it may include both the audio signal in the state of reading aloud in unison and the audio signal emitted by the teacher through the speaker. If only sound energy is used as the standard for judgment, it is easy to make a misjudgment.

[0039] In this embodiment, in order to accurately determine whether the audio signal includes the audio signal corresponding to the chorus reading source, the recording device obtains a preset angle of arrival of interfering sound; based on the angle of arrival of interfering sound, it excludes the angle of arrival of sound associated with the angle of arrival of interfering sound in the angle of arrival information, and obtains the excluded angle of arrival information, thereby realizing the exclusion of interfering sound sources and reducing the impact of interfering sound sources on the judgment of the chorus reading audio signal.

[0040] S3: If the number of sound arrival angles after exclusion exceeds the preset threshold for the number of first sound sources, the audio signal is determined to be unison reading.

[0041] The first threshold for the number of sound sources is the number of sound sources corresponding to the state of reading aloud in unison. Since the audio signal in the state of reading aloud in unison usually includes audio signals emitted by several student sound sources, the number of sound arrival angles after excluding the sound arrival angle information can be compared with the preset first threshold for the number of sound sources to confirm whether the audio signal includes the audio signal in the state of reading aloud in unison.

[0042] In this embodiment, the recording device obtains the number of sound arrival angles of the excluded sound arrival angle information based on the sound arrival angle information after exclusion. If the number of sound arrival angles of the excluded sound arrival angle information exceeds the preset first sound source number threshold, the audio signal is determined to be unison reading. The preset first sound source number threshold is used to indicate the number of sound sources corresponding to unison reading.

[0043] Specifically, if the number of sound arrival angles after exclusion does not exceed the preset threshold for the number of first sound sources, and it is confirmed that the audio signal is not a chorus reading, that is, it does not include audio signals in the chorus reading state, then the audio signal can be directly subjected to noise reduction processing.

[0044] If the number of sound arrival angles after exclusion exceeds the preset threshold for the number of first sound sources, the audio signal is confirmed to be a choral reading, that is, an audio signal in the choral reading state, and the audio signal is used as the target audio signal without noise reduction processing.

[0045] The embodiments of this application acquire audio signals collected in a classroom; based on the audio signals, obtain the sound arrival angle information and sound energy data of the audio signals, wherein the sound arrival angle information includes several sound arrival angles; if the sound energy data of the audio signals exceeds a preset sound energy threshold, obtain a preset interference sound arrival angle; based on the preset interference sound arrival angle, exclude the interference sound arrival angles in the sound arrival angle information to obtain excluded sound arrival angle information; if the number of sound arrival angles in the excluded sound arrival angle information exceeds a preset first sound source number threshold, determine that the audio signal is a choral reading, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to the choral reading. By calculating the sound arrival angle and the sound energy corresponding to the audio signal, accurate judgment of choral reading audio is achieved, reducing costs.

[0046] In an optional embodiment, the recording device may use a method based on the time difference of sound arrival to obtain several sound arrival angles. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 The flowchart of S1 in the audio determination method provided in one embodiment of this application is shown below, including steps S11 to S14, as follows:

[0047] S11: Collect audio signals from the classroom using at least two audio acquisition devices to obtain the audio signals corresponding to the audio acquisition devices.

[0048] In this embodiment, the recording and broadcasting equipment collects audio signals from the classroom through at least two audio acquisition devices to obtain the audio signals corresponding to the audio acquisition devices.

[0049] Specifically, the user sets up audio acquisition devices, such as microphones, at two preset locations. These audio acquisition devices transmit data with the recording equipment. The two audio acquisition devices acquire corresponding audio signals at their respective preset locations and send them to the recording equipment. The audio signals are as follows:

[0050] x1(t)=s(t-τ1)+n1(t)

[0051] x²(t) = s(t-τ²) + n²(t)

[0052] In the formula, x1(t) is the first audio signal acquired by the first audio acquisition device, s(t-τ1) is the sound source signal included in the first audio signal, and n1(t) is the additive noise signal included in the first audio signal; x2(t) is the second audio signal acquired by the second audio acquisition device, s(t-τ2) is the sound source signal included in the second audio signal, and n2(t) is the additive noise signal included in the second audio signal; t is a time parameter, and τ1 and τ2 are the time delay parameters of the first audio acquisition device and the second audio acquisition device, respectively, which are used to reflect the delay time of the audio signal sent by the sound source to reach the first audio acquisition device and the second audio acquisition device.

[0053] S12: Perform Fourier calculation on the audio signal corresponding to the audio acquisition device to obtain the spectrum signal of the audio signal corresponding to the audio acquisition device. Based on the spectrum signal of the audio signal corresponding to the audio acquisition device and the preset cross-spectrum calculation algorithm, obtain the cross-spectrum coefficients of the audio signal.

[0054] The cross-spectrum calculation algorithm is as follows:

[0055]

[0056] In the formula, X1(ω) is the cross-spectral coefficient, X2(ω) is the spectrum of the first audio signal, and X2(ω) is the spectrum of the second audio signal.

[0057] In this embodiment, the recording and broadcasting device performs Fourier calculations on the audio signals corresponding to the audio acquisition device to obtain the spectrum signal of the audio signals corresponding to the audio acquisition device. Based on the spectrum signal of the audio signals corresponding to the audio acquisition device and the preset cross-spectral calculation algorithm, the cross-spectral coefficients corresponding to the audio signals are obtained.

[0058] S13: Calculate the cross-correlation coefficients of the audio signal based on the cross-spectral coefficients of the audio signal, and obtain the set of time delay parameters by taking the set of the cross-correlation coefficients of the audio signal.

[0059] The delay parameter is used to indicate the time difference between the arrival of an audio signal from the same audio source at at least two audio acquisition devices.

[0060] The recording and broadcasting equipment calculates the cross-correlation coefficients corresponding to the audio signal based on the cross-spectral coefficients corresponding to the audio signal, and obtains the set of time delay parameters by taking the set of the cross-correlation coefficients corresponding to the audio signal.

[0061] Specifically, the recording equipment calculates the cross-correlation coefficient of the audio signal based on the cross-spectral coefficients of the audio signal.

[0062] In an optional embodiment, the recording device obtains the cross-correlation coefficient of the audio signal based on the cross-spectral coefficients corresponding to the audio signal and a preset frequency domain weighting function, as follows:

[0063]

[0064] In the formula, Let Ψ be the cross-correlation coefficient. 12 For frequency domain weighting functions, Cross-spectral coefficients;

[0065] Among them, the frequency domain weighting function can be a cross-correlation function, a smooth coherence transformation function, a Roth processing function, a maximum likelihood weighting function, or a PHAT weighting function.

[0066] In another optional embodiment, the recording device obtains the cross-correlation coefficient of the audio signal based on the cross-spectral coefficients corresponding to the audio signal and the entropy value of the absolute value of the cross-spectral coefficients corresponding to the audio signal, as follows:

[0067]

[0068] The recording and broadcasting equipment calculates the set of cross-correlation coefficients corresponding to the audio signals according to a preset set function, obtaining several time delay parameters that maximize the cross-correlation coefficients. These parameters are then combined to obtain a set of time delay parameters. Each time delay parameter represents the time delay between the audio signal sent by the corresponding sound source and the audio acquisition device. The set function is as follows:

[0069]

[0070] S14: Based on the distance data between audio acquisition devices, the set of time delay parameters, and the preset sound arrival angle calculation algorithm, obtain the sound arrival angle corresponding to several time delay parameters.

[0071] The algorithm for calculating the angle of arrival of sound is as follows:

[0072]

[0073] In the formula, L represents the distance, θ represents the angle of sound arrival, and v represents the speed of sound propagation. This is the time delay parameter.

[0074] In this embodiment, the recording and broadcasting device obtains the sound arrival angles corresponding to several time delay parameters based on the distance data between audio acquisition devices, the set of time delay parameters, and the preset sound arrival angle calculation algorithm.

[0075] Please see Figure 3 , Figure 3The flowchart of S2 in the audio determination method provided in one embodiment of this application is shown, which also includes steps S21 to S22, as follows:

[0076] S21: Collect audio signals emitted by interfering sound sources in the classroom multiple times, obtain the audio signal of each collection, and obtain several sound arrival angles corresponding to each collection audio signal.

[0077] The audio signals emitted by the interfering sound sources include the audio signals played out by the loudspeaker and the environmental interference signals. Since teachers often speak through loudspeakers in the classroom, the sound energy corresponding to the audio signals emitted by the loudspeakers is much greater than the energy of normal speech. There may also be environmental interference signals corresponding to noise sources in the surrounding environment.

[0078] To improve the efficiency and accuracy of recording equipment in judging audio signals, the recording equipment can collect audio signals emitted by interfering sound sources in the classroom multiple times using audio acquisition devices to obtain the arrival angle of the interfering sound.

[0079] In order to quickly and accurately obtain the preset angle of arrival of the interference sound, in this embodiment, the recording device collects the audio signal emitted by the interference sound source in the classroom multiple times according to the preset number of samplings, obtains the audio signal collected each time, and obtains several sound arrival angles corresponding to each collected audio signal.

[0080] S22: If the number of several sound arrival angles corresponding to each acquired audio signal exceeds the preset threshold for the number of second sound sources, the same sound arrival angle is obtained based on the several sound arrival angles corresponding to each acquired audio signal, and used as the arrival angle of the interference sound.

[0081] A preset second sound source number threshold is used to indicate the number of sound sources corresponding to interference sound sources. If the number of several sound arrival angles corresponding to each acquired audio signal exceeds the preset second sound source number threshold, then the audio signal includes the audio signal corresponding to the interference sound source. Since the position of the interference sound source, i.e., the speaker sound source, is fixed, the recording and broadcasting equipment analyzes several sound arrival angles corresponding to each acquired audio signal based on the sound arrival angle, obtaining the same sound arrival angle as the interference sound arrival angle. Through multiple detections, the interference sound arrival angle is obtained quickly and accurately.

[0082] For details, please refer to Figure 4 , Figure 4 The flowchart of S2 in the audio determination method provided in one embodiment of this application is shown below, including steps S23 to S24, as follows:

[0083] S23: Construct an interference sound arrival angle range based on the interference sound arrival angle and a preset interference sound arrival angle threshold.

[0084] In order to minimize the influence of interfering sound sources, in this embodiment, the recording and broadcasting device obtains the minimum and maximum interfering sound arrival angles based on the arrival angle of the interfering sound and a preset arrival angle of the interfering sound, and constructs an interfering sound arrival angle range.

[0085] Specifically, the recording equipment obtains the pre-input threshold angle of arrival for interfering sounds, subtracts the threshold angle from the pre-input angle of arrival for interfering sounds, and uses the subtraction result as the minimum angle of arrival for interfering sounds. Then, it adds the pre-input angle of arrival for interfering sounds to the threshold angle of arrival for interfering sounds, and uses the sum as the maximum angle of arrival for interfering sounds, thus constructing an interval of interfering sound arrival angles. The threshold angle of arrival for interfering sounds can be set according to actual conditions and is not limited thereto.

[0086] S24: Determine whether each sound arrival angle in the sound arrival angle information is within the range of interfering sound arrival angles, mark the sound arrival angles within the range of interfering sound arrival angles as interfering sound arrival angles, exclude the interfering sound arrival angles in the sound arrival angle information, and obtain the excluded sound arrival angle information.

[0087] In this embodiment, the recording device determines whether each sound arrival angle in the sound arrival angle information is within the interference sound arrival angle range, marks the sound arrival angles within the interference sound arrival angle range as interference sound arrival angles, excludes the interference sound arrival angles in the sound arrival angle information, and obtains the excluded sound arrival angle information.

[0088] By calculating the sound arrival angle corresponding to the audio signal, several sound arrival angles in the sound arrival angle information of the audio signal collected from the classroom are marked, which realizes the accurate distinction between the arrival angle of interference sound and the arrival angle of normal sound, and reduces costs.

[0089] Please refer to Figure 5 , Figure 5 This is a schematic diagram of an audio judgment device provided in one embodiment of this application. The device can implement all or part of the audio judgment method through software, hardware, or a combination of both. The audio judgment device 5 includes:

[0090] The data acquisition module 51 is used to acquire audio signals collected in the classroom; based on the audio signals, it obtains the sound arrival angle information and sound energy data of the audio signals, wherein the sound arrival angle information includes several sound arrival angles;

[0091] The sound arrival angle processing module 52 is used to obtain a preset interference sound arrival angle if the sound energy data of the audio signal exceeds a preset sound energy threshold; and to exclude the interference sound arrival angle in the sound arrival angle information according to the preset interference sound arrival angle to obtain the excluded sound arrival angle information.

[0092] The target audio signal determination module 53 is used to determine that the audio signal is unison reading if the number of sound arrival angles after exclusion exceeds a preset first sound source number threshold. The preset first sound source number threshold is used to indicate the number of sound sources corresponding to unison reading.

[0093] In an optional implementation, the sound source marking module 51 includes:

[0094] The first audio signal acquisition module is used to acquire audio signals in the classroom through at least two audio acquisition devices to obtain the audio signals corresponding to the audio acquisition devices;

[0095] The cross-spectrum coefficient calculation module is used to perform Fourier calculation on the audio signal corresponding to the audio acquisition device to obtain the spectrum signal of the audio signal corresponding to the audio acquisition device. Based on the spectrum signal of the audio signal corresponding to the audio acquisition device and the preset cross-spectrum calculation algorithm, the cross-spectrum coefficients of the audio signal are obtained.

[0096] The time delay parameter calculation module calculates the cross-correlation coefficients corresponding to the audio signal based on the cross-spectral coefficients corresponding to the audio signal, and obtains a set of time delay parameters by taking the set of the cross-correlation coefficients corresponding to the audio signal. The time delay parameter set includes several time delay parameters, which are used to indicate the time difference between the audio signal sent by the same sound source and at least two audio acquisition devices.

[0097] The sound arrival angle calculation module is used to obtain the sound arrival angle corresponding to several time delay parameters based on the distance data between audio acquisition devices, the set of time delay parameters, and the preset sound arrival angle calculation algorithm.

[0098] In an optional implementation, the sound arrival angle processing module 52 includes:

[0099] The second audio signal acquisition module is used to acquire audio signals emitted by interfering sound sources in the classroom multiple times, obtain the audio signal acquired each time, and obtain several sound arrival angles corresponding to each acquired audio signal.

[0100] The interference sound arrival angle calculation module is used to obtain the same sound arrival angle as the interference sound arrival angle if the number of sound arrival angles corresponding to each acquired audio signal exceeds the preset second sound source number threshold.

[0101] In this embodiment, an audio signal collected in the classroom is acquired through a data acquisition module. Based on the audio signal, the sound arrival angle information and sound energy data of the audio signal are obtained, wherein the sound arrival angle information includes several sound arrival angles. Through a sound arrival angle processing module, if the sound energy data of the audio signal exceeds a preset sound energy threshold, a preset interference sound arrival angle is acquired. Based on the preset interference sound arrival angle, the interference sound arrival angles in the sound arrival angle information are excluded to obtain the excluded sound arrival angle information. Through a target audio signal determination module, if the number of sound arrival angles in the excluded sound arrival angle information exceeds a preset first sound source number threshold, the audio signal is determined to be unison reading, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to unison reading. By calculating the sound arrival angle and the sound energy corresponding to the audio signal, accurate judgment of unison reading audio is achieved, reducing costs.

[0102] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. This application also provides an electronic device, including: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor; the electronic device can store multiple instructions, which are suitable for being loaded and executed by the processor. Figures 1 to 4 The method steps of the illustrated embodiment can be found in the following documentation for detailed execution. Figures 1 to 4 The specific details of the illustrated embodiments will not be elaborated here.

[0103] The processor may include one or more processing cores. The processor 61 connects to various parts within the electronic device using various interfaces and lines. It executes various functions and processes data of the audio judgment device 5 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 62, and by calling data stored in the memory 62. Optionally, the processor 61 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 61 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 61 and may be implemented as a separate chip.

[0104] The memory 62 may include random access memory (RAM) or read-only memory. Optionally, the memory 62 may include a non-transitory computer-readable storage medium. The memory 62 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 62 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch instructions), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data involved in the various method embodiments described above, etc. Optionally, the memory 62 may also be at least one storage device located remotely from the aforementioned processor 61.

[0105] This application embodiment also provides a storage medium that can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figures 1 to 4 For the method steps and specific execution process, please refer to [link / reference]. Figures 1 to 4 The specific details of the illustrated embodiments will not be elaborated here.

[0106] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0107] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0108] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0109] In the embodiments provided in this application, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0111] This application is not limited to the above-described embodiments. If any modifications or variations to this application do not depart from the spirit and scope of this application, and if such modifications and variations fall within the scope of the claims and equivalent technologies of this application, then this application also intends to include such modifications and variations.

Claims

1. An audio judgment method, characterized in that, Includes the following steps: Acquire audio signals collected in the classroom; based on the audio signals, obtain the sound arrival angle information and sound energy data of the audio signals, wherein the sound arrival angle information includes several sound arrival angles; If the sound energy data of the audio signal exceeds a preset sound energy threshold, a preset interference sound arrival angle is obtained; based on the preset interference sound arrival angle, the interference sound arrival angle in the sound arrival angle information is excluded to obtain the excluded sound arrival angle information, so as to exclude interference sound sources in the audio signal and reduce misjudgment of interference sound sources aligning with the reading. If the number of sound arrival angles after exclusion exceeds a preset first sound source number threshold, the audio signal is determined to be unison reading, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to unison reading.

2. The audio judgment method according to claim 1, characterized in that, The process of obtaining the preset angle of arrival of the interference sound includes the following steps: The audio signal emitted by the interfering sound source in the classroom is collected multiple times to obtain the audio signal of each collection and to obtain several sound arrival angles corresponding to each collection audio signal. If the number of several sound arrival angles corresponding to each acquired audio signal exceeds a preset second sound source number threshold, the same sound arrival angle is obtained based on the several sound arrival angles corresponding to each acquired audio signal, and is used as the interference sound arrival angle, wherein the preset second sound source number threshold is used to indicate the number of sound sources corresponding to the interference sound source.

3. The audio judgment method according to claim 2, characterized in that: The audio signals emitted by the interference source include audio signals played from a speaker and environmental interference signals.

4. The audio judgment method according to claim 1, characterized in that, The step of eliminating sound arrival angles associated with the interfering sound arrival angle from the sound arrival angle information based on the interfering sound arrival angle, and obtaining the eliminated sound arrival angle information, includes the following steps: Based on the arrival angle of the interference sound and the preset threshold for the arrival angle of the interference sound, an interval for the arrival angle of the interference sound is constructed. Determine whether each of the sound arrival angles in the sound arrival angle information is within the range of the interfering sound arrival angles, mark the sound arrival angles within the range of the interfering sound arrival angles as interfering sound arrival angles, exclude the interfering sound arrival angles in the sound arrival angle information, and obtain the excluded sound arrival angle information.

5. The audio judgment method according to any one of claims 1 to 4, characterized in that, The process of obtaining audio signals collected in the classroom, and obtaining the sound arrival angle information and sound energy data of the audio signals based on the audio signals, includes the following steps: The audio signal in the classroom is collected by at least two audio acquisition devices to obtain the audio signal corresponding to the audio acquisition device; Fourier transform is performed on the audio signal corresponding to the audio acquisition device to obtain the spectrum signal of the audio signal corresponding to the audio acquisition device. Based on the spectrum signal of the audio signal corresponding to the audio acquisition device and the preset cross-spectral calculation algorithm, the cross-spectral coefficients corresponding to the audio signal are obtained. Based on the cross-spectral coefficients corresponding to the audio signal, the cross-correlation coefficients corresponding to the audio signal are calculated, and a set of cross-correlation coefficients corresponding to the audio signal is obtained to obtain a set of time delay parameters. The set of time delay parameters includes several time delay parameters, which are used to indicate the time difference between audio signals sent from the same sound source and the at least two audio acquisition devices. Based on the distance data between the audio acquisition devices and the set of time delay parameters, the sound arrival angles corresponding to the plurality of time delay parameters are obtained.

6. The audio judgment method according to claim 5, characterized in that, The step of calculating the cross-correlation coefficient of the audio signal based on the cross-spectral coefficients of the audio signal includes the following steps: The cross-correlation coefficients of the audio signal are obtained based on the cross-spectral coefficients corresponding to the audio signal and the preset frequency domain weighting function.

7. The audio judgment method according to claim 5, characterized in that, The step of calculating the cross-correlation coefficient of the audio signal based on the cross-spectral coefficients of the audio signal includes the following steps: The cross-correlation coefficients corresponding to the audio signal are obtained based on the cross-spectral coefficients corresponding to the audio signal and the entropy value of the absolute value of the cross-spectral coefficients corresponding to the audio signal.

8. An audio judging device, characterized in that, include: The data acquisition module is used to acquire audio signals collected in the classroom; Based on the audio signal, obtain the sound arrival angle information and the sound energy data of the audio signal, wherein the sound arrival angle information includes several sound arrival angles; The sound arrival angle processing module is used to obtain a preset interference sound arrival angle if the sound energy data of the audio signal exceeds a preset sound energy threshold; based on the preset interference sound arrival angle, exclude the interference sound arrival angle in the sound arrival angle information to obtain the excluded sound arrival angle information, so as to exclude interference sound sources in the audio signal and reduce misjudgment of interference sound sources aligning with sound reading. The target audio signal determination module is used to determine that the audio signal is a choral reading if the number of sound arrival angles of the excluded sound arrival angle information exceeds a preset first sound source number threshold, wherein the preset first sound source number threshold is used to indicate the number of sound sources corresponding to the choral reading.

9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor; the computer program, when executed by the processor, implements the steps of the audio determination method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the audio judgment method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for automatically muting audios of multiple student terminals

    CN110392301A

  • Audio signal processing method and device, electronic equipment and storage medium

    CN115038014A