A listening training system for English teaching

By recording BRIR raster library and noise template library in the examination room, and combining triangular element centroid interpolation and volume adjustment, virtual reflected sound data is generated, which solves the problems of sound pressure level consistency and insufficient noise simulation in traditional English listening training systems, and realizes a realistic examination room listening experience and anti-interference training.

CN120954287BActive Publication Date: 2025-12-23SICHUAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511460340.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-12-23
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Traditional English listening training systems cannot realistically simulate complex acoustic environments, resulting in insufficient consistency of sound pressure levels, limited noise simulation scenarios, and a lack of interactivity and flexibility, making it difficult to meet the needs of targeted training of specific seat listening skills.

Method used

By uniformly arranging reference points in the examination room to record BRIR grid libraries and noise template libraries, and combining triangular element centroid interpolation and volume adjustment parameters, virtual reflected sound data is generated, and the listening training signal of the target seat is synthesized to simulate the acoustic environment in the examination room.

Benefits of technology

It achieves a highly realistic exam room listening experience through binaural headphones, improves students' adaptability and anti-interference ability to complex exam sound fields, and solves the problem of inconsistency between traditional training and the real environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954287B_ABST
    Figure CN120954287B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of hearing training, and discloses an English teaching hearing training system, G*H reference points are arranged in a row-column grid in a target examination room, left and right ear sound reflection data are recorded at the G*H reference points, and the left and right ear sound reflection data are stored as a BRIR grid library; meanwhile, examination room natural noise is recorded at the G*H reference points, and the examination room natural noise is saved as a noise template library; broadcasting content is recorded at adjacent examination rooms of the target examination room, and the broadcasting content is saved as adjacent room voice files; reference sound pressure levels are measured at a preset distance from an examination room loudspeaker; target seat coordinates set by a user are determined; three reference points corresponding to the target seat coordinates are determined in the BRIR grid library, left and right ear sound reflection data of the three reference points are mixed in a triangle element formed by the three reference points based on the target seat coordinates according to weights, and virtual reflection sound data of the target seat coordinates is obtained; the application can simulate the acoustic environment of any seat in an examination room, can be used for self-selected position training, and can improve hearing adaptability and anti-interference capability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of hearing training, more particularly, it relates to a hearing training system for English teaching. BACKGROUND

[0002] In the field of English teaching, hearing training is a key link to improve language comprehensive ability, but the traditional training mode faces significant challenges. The current mainstream hearing teaching mainly relies on fixed recording materials playing in standard acoustic environment, which is difficult to simulate the complex acoustic environment variables in real scenes, such as spatial position difference, background noise interference, cross-region sound transmission, etc. For students who need to cope with standardized tests, the acoustic conditions of different seats may significantly affect the hearing effect, and the traditional training cannot let students adapt to these differences in advance.

[0003] In the prior art, part of the hearing training system tries to introduce virtual acoustic simulation technology. For example, the sound field synthesis method based on binaural room impulse response can convert a single-channel signal into a binaural signal through convolution operation, and simulate the sound reflection characteristics of a specific space.

[0004] However, such technology has the following limitations:

[0005] 1. Insufficient sound pressure level consistency, the existing scheme does not establish a physical mapping model of sound pressure level and spatial distance, so that the loudness of the sound played by the earphone cannot truly reflect the sound pressure attenuation law of the actual seat, causing the hearing to be disconnected with the real experience;

[0006] 2. Single noise simulation scene, most systems can only superimpose fixed types of background noise, and cannot simulate the composite interference source in the examination room, which contains both the environmental noise of the scene and the speech transmission of the adjacent scene, and it is difficult to reproduce the multi-source acoustic interference in the real examination;

[0007] 3. Lack of interaction and flexibility, the existing system cannot support students to actively select seat coordinates, and it is difficult to meet the needs of targeted training of specific seat hearing. SUMMARY

[0008] The present application provides a hearing training system for English teaching, which solves the technical problems raised in the background art.

[0009] The present application provides a hearing training system for English teaching, which includes:

[0010] The offline module includes:

[0011] Step 1, uniformly arrange GxH reference points in the target examination room according to the row and column grid, record the left and right ear sound reflection data at each reference point respectively, and store it as a BRIR grid library;

[0012] Step 2, when step 1 is performed, the following operations are performed synchronously:

[0013] Record the natural noise of the test room at each reference point and save it as a noise template library;

[0014] Record the broadcasting content in the adjacent test room of the target test room and save it as an adjacent room audio file;

[0015] Step 3, measure the reference sound pressure level at a preset distance from the test room loudspeaker;

[0016] The training module comprises:

[0017] Step 4, determine the target seat coordinates set by the user;

[0018] Step 5, determine the three reference points corresponding to the target seat coordinates and the corresponding left and right ear sound reflection data in the BRIR grid library, and fuse to obtain the virtual reflection sound data of the target seat coordinates;

[0019] Step 6, determine the volume adjustment parameter according to the Euclidean distance between the target seat coordinates and the test room loudspeaker and the reference sound pressure level;

[0020] Step 7, convolve the hearing training audio with the virtual reflection sound data and apply the volume adjustment parameter to generate left and right ear clean sound signals;

[0021] Step 8, select the natural noise of the reference point with the smallest Euclidean distance from the target seat coordinates from the noise template library, and determine the noise adjustment parameter according to the preset target signal-to-noise ratio to generate an adjusted noise;

[0022] Obtain the adjacent test room broadcasting from the adjacent room audio file, and determine the adjacent test room broadcasting adjustment parameter according to the preset wall attenuation amount to generate an adjusted adjacent test room broadcasting;

[0023] Mix the adjusted noise and the adjusted adjacent test room broadcasting into the left and right ear clean sound signals to obtain the target left and right ear sound signals of the target seat coordinates.

[0024] Further, GxH reference points are uniformly arranged in the target test room according to the row and column grid, left and right ear sound reflection data are recorded at the GxH reference points respectively, and stored as a BRIR grid library, comprising:

[0025] Measure the horizontal width of the target test room and the longitudinal length , and set the origin of the site coordinate system ;

[0026] Based on the horizontal width , the longitudinal length , and the grid row and column number G and H, respectively calculate the horizontal spacing and longitudinal spacing ;

[0027] Load preset test signal , This represents the signal at the nth sampling point in the preset test signal. And it will be played through the loudspeakers in the examination room;

[0028] In coordinates Record the left and right ear response signals separately. and ;

[0029] Response signals to left and right ears and With test signal Perform deconvolution operation:

[0030] ,

[0031] ,

[0032] in, This represents the sound reflection data from the left ear. This represents the sound reflection data from the right ear. Represents the Discrete Fourier Transform. Indicates the inverse discrete Fourier transform;

[0033] Will and Organized and stored as a BRIR raster library by row and column indexes.

[0034] Furthermore, natural noise in the examination room is recorded at G×H reference points and saved as a noise template library; broadcast content is recorded in adjacent examination rooms of the target examination room and saved as adjacent field speech files; including:

[0035] Record natural noise signals within the target examination room at G×H reference points for a preset duration. , This represents the signal at the nth sampling point in the natural noise signal, where 1 ≤ n ≤ N;

[0036] For natural noise signals Calculate the sound pressure level at the corresponding reference point ,as follows:

[0037]

[0038] in, This indicates the number of sampling points within a preset time period;

[0039] Record a pre-set duration audio signal in an examination room adjacent to the target examination room. , represents the n-th sample point signal in the voice content signal;

[0040] the voice content signal calculating the sound pressure level as follows:

[0041] ,

[0042] the natural noise signal and the corresponding sound pressure level are stored as a noise template library;

[0043] the voice content signal and the corresponding sound pressure level are stored as a neighbor field voice file.

[0044] Further, three reference points corresponding to the target seat coordinates and the corresponding left and right ear sound reflection data in the BRIR grid library are determined, and the virtual reflection sound data of the target seat coordinates are fused, including:

[0045] determining the target seat coordinates P set by the user, ;

[0046] By GxH reference points, the BRIR grid library is discretized into a plurality of triangular elements;

[0047] determining the triangular element containing the target seat coordinates, obtaining the reference points A, B and C corresponding to the triangular element;

[0048] respectively determining the coordinates of the reference points A, B and C as , , ;

[0049] calculating the triangular element barycentric coordinate weight:

[0050] ,

[0051] ,

[0052] ,

[0053]

[0054] wherein, represents the triangular element barycentric coordinate weight normalization factor, , and , respectively, represent the linear fusion weight of the reference points A, B and C;

[0055] Take the sound reflection data of the left and right ears corresponding to reference points A, B, and C respectively, and weight them according to the centroid coordinates of the triangular elements. , and Linear fusion is performed to obtain virtual reflected sound data of the target seat.

[0056] Furthermore, based on the target seat coordinates and the Euclidean distance between the target seat and the examination room loudspeakers, as well as the reference sound pressure level, the volume adjustment parameters are determined, including:

[0057] Based on the target seat coordinates And the pre-recorded coordinates of the exam room speakers ;

[0058] Determine the Euclidean distance between the target seat coordinates and the examination room speaker coordinates. ;

[0059] Determine unit distance based on prior knowledge Corresponding reference sound pressure level ;

[0060] Based on reference distance and reference sound pressure level Calculate the target sound pressure level using the free-field attenuation model:

[0061]

[0062] in, Indicates the target sound pressure level;

[0063] Determine the volume adjustment parameters as follows:

[0064]

[0065] in, This indicates the volume adjustment parameters.

[0066] Furthermore, the audio data from the hearing training is convolved with the virtual reflected sound data, and volume adjustment parameters are applied to generate clean sound signals for both ears, including:

[0067] Get listening training audio Virtual reflected sound data of the target seat and and volume adjustment parameters ;

[0068] The virtual reflected sound data and the audio used for listening training are subjected to discrete convolution operations as follows:

[0069] ,

[0070] ,

[0071] in, express The amplitude of the intermediate signal after convolution at the nth sampling point. express In the The amplitude of each sampling point This indicates that the audio for listening training is in the [number]th [section]. The amplitude of each sampling point Indicates a time delay. express The amplitude of the intermediate signal after convolution at the nth sampling point. express In the The amplitude of each sampling point Let Q-1 represent the number of sampling points in the listening training audio. This represents the convolution operation;

[0072] Will and Respectively with volume adjustment parameters Multiply them to generate clean sound signals for both ears.

[0073] Furthermore, the natural noise of the reference point with the smallest Euclidean distance to the target seat coordinates is selected from the noise template library, and the noise adjustment parameters are determined according to the preset target signal-to-noise ratio to generate the adjusted noise, including:

[0074] Calculate the Euclidean distance between each reference point in the noise template library and the target seat coordinates. ;

[0075] Determine The smallest reference point, and the natural noise signal corresponding to the reference point is selected.

[0076] Based on the clean sound signal from the left ear Signal-to-noise ratio compared to preset target Calculate noise adjustment parameters ,as follows:

[0077]

[0078] in, This indicates the calculation of the Euclidean norm;

[0079] natural noise signal With noise adjustment parameters Multiply to generate adjusted noise.

[0080] Further, the adjacent test room voice file is obtained to acquire an adjacent test room voice, and an adjacent test room voice adjustment parameter is determined according to a preset wall signal attenuation amount, and the adjusted adjacent test room voice is generated, including:

[0081] obtaining a voice content signal from the adjacent test room voice file ;

[0082] obtaining a wall signal attenuation amount of the target test room and the adjacent test room ;

[0083] converting the wall signal attenuation amount into a voice adjustment parameter ; , ;

[0084] multiplying the voice content signal and the voice adjustment parameter to generate the adjusted adjacent test room voice.

[0085] Further, the adjusted noise and the adjusted adjacent test room voice are mixed into the left and right ear clean sound signals to obtain the target left and right ear sound signals of the target seat coordinate, including:

[0086] determining a left ear clean sound signal , a right ear clean sound signal , the adjusted noise and the adjusted adjacent test room voice ;

[0087] calculating the target left ear sound signal of the target seat coordinate, as follows:

[0088]

[0089] wherein, represents the target left ear sound signal of the target seat coordinate;

[0090] calculating the target right ear sound signal of the target seat coordinate, as follows:

[0091]

[0092] wherein, represents the target left ear sound signal of the target seat coordinate.

[0093] The beneficial effect of the present application is that the listening experience of any seat in the examination room can be highly realistically restored in the binaural earphone environment, by constructing a multi-dimensional acoustic database containing sound source reflection, on-site noise, adjacent interference and sound pressure level change, and combining real-time interpolation and signal synthesis technology, so that students can actively choose different acoustic environments of the space position in the training, thereby significantly improving their adaptability and anti-interference ability to complex examination sound field, effectively solving the problem that the traditional earphone hearing training and the real hearing test environment are inconsistent. BRIEF DESCRIPTION OF DRAWINGS

[0094] Figure 1 is a module diagram of an English teaching listening training system of the present application. DETAILED DESCRIPTION

[0095] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is merely meant to provide a better understanding of the subject matter described herein and can be changed in function and arrangement without departing from the scope of the present description. Various examples can omit, substitute, or add various procedures or components as appropriate or desired. In addition, features described in relation to some examples can also be combined in other examples.

[0096] As shown in Figure 1 , an English teaching listening training system comprises:

[0097] An offline module comprises:

[0098] Step 1, GxH reference points are arranged in a row-column grid in the target examination room, left and right ear sound reflection data are recorded at each reference point respectively, and stored as a BRIR grid library;

[0099] Step 2, when step 1 is executed, the following operations are performed synchronously:

[0100] Record the natural noise of the examination room at each reference point m, and save it as a noise template library;

[0101] Record the broadcast content in the adjacent examination room of the target examination room, and save it as an adjacent room speech file;

[0102] Step 3, measure the reference sound pressure level at a predetermined distance from the examination room loudspeaker;

[0103] A training module comprises:

[0104] Step 4, determine the target seat coordinates set by the user;

[0105] Step 5, determining three reference points corresponding to the target seat coordinate and the corresponding left and right ear sound reflection data in the BRIR grid library, and fusing to obtain the virtual reflection sound data of the target seat coordinate;

[0106] Step 6, determining the volume adjustment parameter according to the Euclidean distance between the target seat coordinate and the test room loudspeaker and the reference sound pressure level;

[0107] Step 7, convolving the hearing training audio with the virtual reflection sound data, and applying the volume adjustment parameter to generate left and right ear clean sound signals;

[0108] Step 8, selecting the natural noise of the reference point with the smallest Euclidean distance from the target seat coordinate from the noise template library, and determining the noise adjustment parameter according to the preset target signal-to-noise ratio to generate the adjusted noise;

[0109] Obtaining the adjacent test room broadcast from the adjacent test room voice file, and determining the adjacent test room broadcast adjustment parameter according to the preset wall attenuation amount to generate the adjusted adjacent test room broadcast;

[0110] Mixing the adjusted noise and the adjusted adjacent test room broadcast into the left and right ear clean sound signals to obtain the target left and right ear sound signals of the target seat coordinate.

[0111] In an embodiment of the present application, GxH reference points are uniformly arranged in rows and columns in the target test room, left and right ear sound reflection data are recorded at the GxH reference points respectively, and the left and right ear sound reflection data are stored as a BRIR grid library, including:

[0112] Measuring the horizontal width and the longitudinal length of the target test room, and setting the origin of the test room coordinate system ;

[0113] Based on the horizontal width , the longitudinal length , and the grid row and column number G and H, the horizontal spacing and the longitudinal spacing are calculated respectively;

[0114] Loading a preset test signal , , wherein n represents the nth sampling point signal in the preset test signal, , and playing through the test room loudspeaker;

[0115] Recording left and right ear response signals and at coordinates respectively;

[0116] The left and right ear response signals and are convolved with the test signal A deconvolution operation is performed:

[0117]

[0118]

[0119] wherein, represents left ear sound reflection data, represents right ear sound reflection data, represents a discrete Fourier transform, represents an inverse discrete Fourier transform;

[0120] are stored as a BRIR grid library. are stored as a BRIR grid library.

[0121] In this embodiment, in order to generate an accurate room binaural impulse response (BRIR) for any seat coordinate in the subsequent online phase, it is necessary to first grid sample the target examination room and measure the left and right ear reflection characteristics of each sampling point. The specific operation is as follows:

[0122] The examination room plane is represented by a two-dimensional rectangular coordinate system, and the origin is selected as a corner of the site, with coordinates . The physical dimensions of the examination room in the horizontal direction (width) and the vertical direction (depth) are X meters and Y meters, respectively. According to the experimental design, the number of grid rows is G and the number of grid columns is H, where G≥2 and H≥2.

[0123] First, the equal interval spacing of the grid in the horizontal and vertical directions is calculated:

[0124] The horizontal spacing and the vertical spacing ; in the above formula, the denominators and are used to ensure that the grid endpoints can cover the examination room boundaries, and when i takes the value 1, it takes the value , and when i takes the value G, it takes the value ; similarly in the vertical direction. The parameters and respectively represent the horizontal and vertical distances between adjacent two rows and two columns of reference points, which establish the uniform distribution positions of the G×H sampling points.

[0125] According to the above spacing, the dummy head measurement system is placed at the coordinates , , in turn. The system includes a high-fidelity recording head to simulate the left and right ear positions of a human ear, and records the left and right ear response signals of a known test signal played by the main examination room loudspeaker as and​ .

[0126] To obtain the impulse response data from the above response, it is necessary to perform a deconvolution operation on and the known test signal Here, the method of frequency domain deconvolution is adopted:

[0127]

[0128]

[0129] wherein, represents a discrete Fourier transform, used to map a time domain signal to a frequency domain; represents an inverse discrete Fourier transform, used to map a frequency domain ratio back to a time domain;

[0130] The numerator represents the spectrum of the recorded left ear response, and the denominator is the spectrum of the input test signal, and the division can remove the input characteristics and retain the room reflection characteristics.

[0131] The obtained and are the left and right ear room impulse responses at the reference point .

[0132] After the above calculation, all and are organized according to the row and column indexes to construct a BRIR grid library with a size of GxH. The BRIR grid library contains both the binaural impulse responses of each spatial sampling position and the reverberation, reflection and diffraction characteristics of the room.

[0133] With the BRIR grid library, interpolation algorithms can be used to expand the discrete sampling to any coordinates, realizing high-precision and multi-position spatial acoustic simulation.

[0134] In an embodiment of the present application, the natural noise in the examination room is recorded at GxH reference points and saved as a noise template library; the broadcast content is recorded in the adjacent examination room of the target examination room and saved as a neighbor speech file; including:

[0135] At the GxH reference points in the target examination room, the natural noise signal in the target examination room is recorded for a predetermined duration , represents the signal of the nth sampling point in the natural noise signal, 1≤n≤N;

[0136] The sound pressure level of the corresponding reference point is calculated for the natural noise signal , as follows:

[0137]

[0138] wherein, represents the number of sampling points within a preset time length;

[0139] recording a broadcasting content signal lasting a preset time length in an adjacent examination room to the target examination room , represents the nth sampling point signal in the broadcasting content signal;

[0140] calculating the sound pressure level of the broadcasting content signal as follows:

[0141]

[0142] storing the natural noise signal and the corresponding sound pressure level as a noise template library;

[0143] storing the broadcasting content signal and the corresponding sound pressure level as an adjacent room voice file.

[0144] In the embodiment, in order to accurately synthesize the background noise and adjacent room interference into the listening signal of the target seat, the following is specifically implemented:

[0145] First, G×H noise collection reference points are arranged in the target examination room in an equidistant manner according to a preset number of rows G and a preset number of columns N. For the reference point in the ith row and the jth column, the plane coordinates are denoted as At each reference point, a high-fidelity microphone of the same type as the BRIR collection is used to continuously record an examination room natural noise signal lasting a predetermined time length, denoted as:

[0146] ,

[0147] wherein, represents the total number of sampling points in the recording process, is the nth sampling point signal. In order to quantify the noise energy level of each position, the sound pressure level of the reference point is calculated by using a mean square power to decibel conversion formula:

[0148]

[0149] wherein is the mean square power of the natural noise signal (i.e., the RGS energy);

[0150] ​​For converting power ratio to decibel (dB) scale, facilitating subsequent signal-to-noise ratio calculation and unified calibration.

[0151] Then, the same length of broadcasting content signals are recorded in the examination room adjacent to the target examination room, denoted as:

[0152] , ;

[0153] And the sound pressure level of the broadcasting in the adjacent room is obtained by the same RGS energy calculation method:

[0154] The natural noise signals of each reference point and the corresponding sound pressure level are stored together, that is, to constitute a noise template library; the adjacent room broadcasting signal and the sound pressure level are stored as an adjacent room voice file.

[0155] In an embodiment of the present application, three reference points corresponding to the target seat coordinates and the corresponding left and right ear sound reflection data in the BRIR grid library are determined, and the virtual reflection sound data of the target seat coordinates are fused, including:

[0156] The target seat coordinates P set by the user are determined, ;

[0157] The BRIR grid library is discretized into a plurality of triangular elements by GxH reference points;

[0158] The triangular element containing the target seat coordinates is determined, and the reference points A, B and C corresponding to the triangular element are obtained;

[0159] The coordinates of the reference points A, B and C are respectively determined as , , ;

[0160] The triangular element barycentric coordinate weight is calculated:

[0161]

[0162]

[0163]

[0164]

[0165] wherein, the triangular element barycentric coordinate weight normalization factor is represented by , and respectively represent the linear fusion weight of reference points A, B and C;

[0166] respectively take the left and right ear sound reflection data corresponding to reference points A, B and C, and perform linear fusion according to the triangular element barycentric coordinate weight 、 and to obtain the virtual reflection sound data of the target seat.

[0167] In this embodiment, in order to generate continuous, smooth and physically self-consistent binaural room impulse responses (BRIR) for any target seat coordinates in the online stage, the BRIR of the discrete sampling points is weighted and mixed by using the triangular element barycentric interpolation method on the basis of the BRIR grid library established offline. The specific implementation steps are as follows:

[0168] Firstly, the system receives the target seat plane coordinates specified by the user , and accurately locates the unique triangular element containing point P in the triangular elements obtained after pre-processing the examination room by Delaunay triangulation. The three vertices of the triangular element correspond to the reference points A, B and C in the BRIR grid library, and their coordinates are respectively denoted as:

[0169] 、 、

[0170] Here, the Delaunay triangulation ensures that there is no overlap or excessive narrowness between the vertices of the selected triangular element, thereby improving the stability and physical reasonableness of the interpolation.

[0171] Next, according to the geometric definition of the barycentric coordinates, the weight coefficients of the target point P in the triangular element ABC relative to the vertices A, B and C are calculated. First, the area scalar corresponding to the triangular element ABC is calculated:

[0172]

[0173] wherein is proportional to the actual geometric area of the triangular element, and ensures that the three points A, B and C are not collinear. Then, the following are calculated in turn:

[0174]

[0175]

[0176]

[0177] wherein the weights 、 and All are non-negative and the sum of the three is 1, reflecting the relative position relationship between the target point P and each vertex of the triangular element. This interpolation method is mathematically equivalent to taking the ratio of the areas of the three sub-triangles as weights, thereby achieving linear continuous interpolation of the BRIR at any point in a two-dimensional plane.

[0178] Finally, the left and right ear sound reflection data corresponding to the reference points A, B and C are taken out respectively, and the weighted sum is calculated according to the barycentric weight.

[0179] The barycentric interpolation method ensures that the obtained BRIR is smoothly transitioned in space and does not have abrupt distortion at discrete sampling points, while maintaining the linear superposition characteristics of the physical acoustic response, thereby providing a high-precision spatial filtering kernel for the subsequent convolution synthesis of the original training audio.

[0180] In an embodiment of the present application, the volume adjustment parameter is determined according to the Euclidean distance between the target seat coordinates and the test room speaker, and the reference sound pressure level, including:

[0181] According to the target seat coordinates and the pre-recorded test room speaker coordinates ;

[0182] Determine the Euclidean distance between the target seat coordinates and the test room speaker coordinates ;

[0183] Determine the unit distance corresponding to the reference sound pressure level ;

[0184] Based on the reference distance and the reference sound pressure level , the target sound pressure level is calculated according to the free field attenuation model:

[0185]

[0186] wherein, represents the target sound pressure level;

[0187] Determine the volume adjustment parameter as follows:

[0188]

[0189] wherein, represents the volume adjustment parameter.

[0190] In this embodiment, in order to realize the online stage to present the original training audio at the earphone end with the same sound pressure level as any seat in the real test room, a volume adjustment parameter calculation sub-step based on a spatial distance attenuation model is introduced. Specifically, the steps include:

[0191] First, the system receives the target seat coordinates set by the user:

[0192]

[0193] And the coordinates of the examination room speakers measured and recorded in advance during the offline phase, in the same coordinate system.

[0194]

[0195] This step is used to map the virtual training environment to the actual test room space geometry, ensuring that subsequent sound pressure calculations are based on real spatial relationships.

[0196] Then, calculate the Euclidean distance between the target seat P and the examination room speaker S. ;

[0197] Next, the reference distance during the offline phase is invoked. Reference sound pressure level measured at a distance of 1 meter (for example) (Unit: dB), and based on the ideal attenuation law of free-field sound waves, the sound pressure level varies with distance according to... Proportional attenuation is used to calculate the theoretical sound pressure level at the target seat:

[0198]

[0199] Since there is no additional reflection or absorption in a free field, the relationship between sound pressure level and distance strictly follows the inverse square law.

[0200] Finally, in order to adjust the clean sound signal obtained after convolution to the above sound pressure level, the sound pressure level difference needs to be converted into an amplitude gain coefficient (volume adjustment parameter), which can be obtained according to the definition of decibels:

[0201]

[0202] Among them, molecules This represents the change in sound pressure level relative to a reference distance (positive values ​​indicate gain, negative values ​​indicate attenuation), divided by... The linear amplitude ratio can be obtained by taking the decimal exponent. Amplitude gain coefficient It directly affects each sampling point of the time-domain signal, thereby ensuring that the final playback volume is consistent with the real spatial sound pressure.

[0203] Through the above steps, the sound pressure relationship between the target seat and the speaker is accurately matched in the binaural headphones, so as to achieve a high degree of reproduction of the human ear's perception of the distance of the sound source, and provide students with a listening training experience that corresponds one-to-one with the real examination environment.

[0204] In one embodiment of the present application, the hearing training audio is convolved with the virtual reflected sound data, and a volume adjustment parameter is applied, to generate left and right ear clean sound signals, including:

[0205] Obtaining a hearing training audio , virtual reflected sound data of a target seat , and a volume adjustment parameter ;

[0206] Performing a discrete convolution operation on the virtual reflected sound data and the hearing training audio, as follows:

[0207]

[0208]

[0209] wherein, represents the amplitude of the intermediate signal after the convolution operation at the n th sampling point, represents the amplitude at the th sampling point, represents the amplitude of the hearing training audio at the th sampling point, represents the time delay, represents the amplitude of the intermediate signal after the convolution operation at the n th sampling point, represents the amplitude at the th sampling point, represents Q-1, Q being the number of sampling points in the hearing training audio, represents the convolution operation;

[0210] and are respectively multiplied by the volume adjustment parameter , to generate left and right ear clean sound signals.

[0211] In this embodiment, the online training module first obtains a single-channel hearing training audio from a training material library , virtual reflected sound data of a target seat , and a volume adjustment parameter .

[0212] For time-domain convolution operation, let Q be the total number of sampling points of the original training audio , and the reflection response length is R+1, wherein R=Q−1. For the n th time, the left and right ear intermediate signals are calculated respectively: ​

[0213]

[0214]

[0215] wherein, and respectively represent the amplitude value of the left ear and right ear reflected response at the rth discrete time, represents the amplitude of the training audio at the th sampling time, represents the discrete convolution. The convolution operation superimposes the clean audio signal with the acoustic filter kernel of the target seat, thereby reproducing the room reflection, reverberation and diffraction behavior at the earphone end.

[0216] After completing the convolution, the obtained intermediate signal , still corresponds to the output under the reference sound pressure level . To match the target sound pressure level , the amplitude gain coefficient needs to be applied, which is calculated by the previous distance attenuation model. Thus, the final left and right ear clean sound signals are:

[0217]

[0218]

[0219] wherein the amplitude gain coefficient , if is greater than 1, it is used for amplification, and if is less than 1, it is used for attenuation, ensuring that the convolution synthesized signal reaches the same loudness level as the real examination room target seat coordinates in human ear perception.

[0220] In an embodiment of the present application, the natural noise of the reference point with the smallest Euclidean distance from the target seat coordinates is selected from the noise template library, and the noise adjustment parameter is determined according to the preset target signal-to-noise ratio to generate the adjusted noise, including:

[0221] The Euclidean distance between each reference point in the noise template library and the target seat coordinates is calculated;

[0222] The reference point that makes the smallest is determined, and the natural noise signal corresponding to the reference point is selected

[0223] The noise adjustment parameter is calculated according to the left ear clean sound signal and the preset target signal-to-noise ratio , as follows:

[0224]

[0225] wherein, represents calculating the Euclidean norm;

[0226] The natural noise signal is multiplied by the noise adjustment parameter to generate the adjusted noise.

[0227] When it is necessary to superimpose the examination room natural noise on the clean signal that has completed reflection and volume correction, the system first automatically selects the noise data of the reference point closest to the current target seat from the noise template library generated offline, and adjusts the energy of the noise according to the preset target signal-to-noise ratio, as follows:

[0228] The system knows the plane coordinates of the target seat , and the coordinates of the i-th row and j-th column reference point in the noise template library , and the corresponding noise time domain sequence . First, for each reference point in the noise template library, calculate its Euclidean distance with P .

[0229] The Euclidean distance with P quantifies the spatial separation degree between the target seat coordinates and each reference point, which is used to determine which reference point's noise best represents the local background environment of the target seat coordinates.

[0230] The system selects the reference point that makes the minimum, and takes out the corresponding natural noise signal ;

[0231] Next, the system has previously calculated the left ear clean sound signal, and the target signal-to-noise ratio expected to be achieved in the training task (unit: decibel, self-defined parameter).

[0232] In order to ensure that the natural noise signal after superposition reaches the predetermined value when synthesized with the left ear clean sound signal in the earphone, the natural noise signal needs to be multiplied by a suitable amplitude coefficient .

[0233] This coefficient is given by the following formula:

[0234]

[0235] wherein, represents the Euclidean norm of the left ear clean sound signal, that is, the total energy (proportion of RMS energy) of the left ear clean sound signal;

[0236] the Euclidean norm of the selected reference point noise signal;

[0237] the exponential term then, according to the definition of decibel, the signal-to-noise ratio is converted into a linear amplitude ratio;

[0238] First, by normalizing the natural noise signal to the signal energy, and then multiplying it by to achieve the required signal-to-noise ratio attenuation. The result is a dimensionless amplitude amplification (or attenuation) coefficient, when > 0, < 1, indicating that the adjusted noise will be adjusted down.

[0239] Finally, the system multiplies the natural noise signal by to generate the adjusted noise that meets the target signal-to-noise ratio requirement.

[0240] In an embodiment of the present application, the adjacent test room broadcast is obtained from the adjacent test room voice file, and the adjacent test room broadcast adjustment parameter is determined according to the preset wall signal attenuation, to generate the adjusted adjacent test room broadcast, comprising:

[0241] obtaining the broadcast content signal from the adjacent test room voice file;

[0242] obtaining the wall signal attenuation of the target test room and the adjacent test room;

[0243] converting the wall signal attenuation into a broadcast adjustment parameter , ;

[0244] multiplying the broadcast content signal by the broadcast adjustment parameter to generate the adjusted adjacent test room broadcast.

[0245] In this embodiment, in order to reproduce the real interference volume from the adjacent test room through the wall to the target test room in the earphone, after generating the left and right ear clean signals, the adjacent test room broadcast also needs to be attenuated. Specifically, the system first reads the broadcast content signal from the adjacent test room voice file;

[0246] the broadcast content signal Reflects the original volume and spectral characteristics of the adjacent test room loudspeaker broadcast. At the same time, the system calls the predetermined wall attenuation parameter TL (unit: dB), which is determined based on the acoustic test of the wall material and structure between the target test room and the adjacent test room, and represents the broadcast content signal The energy loss through the wall.

[0247] According to the conversion relationship between decibels and linear amplitude in acoustics, the wall attenuation TL is converted into a linear amplitude attenuation coefficient using the following formula

[0248]

[0249] The greater TL is, the more serious the broadcast content signal energy loss is, and the smaller the corresponding is; when TL is 0 dB, = 1, indicating no attenuation. The conversion is based on the definition of decibels: a 20 dB difference corresponds to an amplitude ratio of 10 times, so the attenuation is taken as negative and divided by 20, and then a decimal exponential operation is performed to obtain the linear amplitude ratio.

[0250] Finally, the N sample point signals in the broadcast content signal are multiplied by sample by sample to obtain the adjusted adjacent test room broadcast.

[0251] In an embodiment of the present application, the adjusted noise and the adjusted adjacent test room broadcast are mixed into the left and right ear clean sound signals to obtain the target left and right ear sound signals of the target seat coordinates, including:

[0252] The left ear clean sound signal , the right ear clean sound signal , the adjusted noise , and the adjusted adjacent test room broadcast are determined.

[0253] The target left ear sound signal of the target seat coordinates is calculated as follows:

[0254]

[0255] Wherein, represents the target left ear sound signal of the target seat coordinates.

[0256] The target right ear sound signal of the target seat coordinates is calculated as follows:

[0257]

[0258] Wherein, represents the target left ear sound signal of the target seat coordinates. ​

[0259] In the present embodiment, the system first acquires the left ear clean sound signal , the right ear clean sound signal , the adjusted noise , and the adjusted adjacent classroom broadcast .

[0260] Each signal is aligned in sampling points and time length, ensuring that the same sampling index n corresponds to the same time point.

[0261] Subsequently, the system adds the clean sound and the noise, interference in turn according to the physical principle of linear superposition of multiple sound sources in the sound field, to generate the left and right ear final play signals of the target seat:

[0262]

[0263]

[0264] Among them, and respectively represent the clean sound sample of the target seat at the left and right ears without any background or interference; represents the natural noise that has been fine-tuned according to the target signal-to-noise ratio; represents the adjacent classroom broadcast that has been corrected according to the wall attenuation TL. Through sample-by-sample addition, the superposition effect of multiple sound sources in the classroom space is simulated, reflecting the final perception of sound at the earphone end.

[0265] The implementation of this mixing step can ensure that the signals , played to the students have both room reflection characteristics and real sound pressure, as well as local noise and adjacent classroom interference, thereby achieving high restoration of the target seat listening experience and anti-interference training requirements.

[0266] The above describes the embodiments of the present embodiment, but the present embodiment is not limited to the specific embodiments described above, which are only illustrative and not limiting. Those skilled in the art can make many forms under the inspiration of the present embodiment, which are all within the protection of the present embodiment.

Claims

1. A listening training system for English teaching, characterized by comprising: Comprise: Offline module, comprising: Step 1, in the target examination room, GxH reference points are arranged in a row-column grid, left and right ear sound reflection data are recorded at each reference point respectively, and stored as a BRIR grid library; Step 2, during the execution of step 1, the following operations are performed synchronously: Record the natural noise of the examination room at each reference point and save it as a noise template library; Record the broadcast content in the adjacent examination room of the target examination room and save it as a neighbor room audio file; Step 3, measure the reference sound pressure level at a predetermined distance from the examination room loudspeaker; Training module, comprising: Step 4, determine the target seat coordinates set by the user; Step 5, determine the three reference points corresponding to the target seat coordinates and the corresponding left and right ear sound reflection data in the BRIR grid library, and fuse to obtain the virtual reflection sound data of the target seat coordinates; Step 6, determine the volume adjustment parameter according to the Euclidean distance between the target seat coordinates and the examination room loudspeaker, and the reference sound pressure level; Step 7, convolve the hearing training audio with the virtual reflection sound data, and apply the volume adjustment parameter to generate left and right ear clean sound signals; Step 8, select the natural noise of the reference point with the smallest Euclidean distance from the target seat coordinates from the noise template library, and determine the noise adjustment parameter according to the preset target signal-to-noise ratio to generate the adjusted noise; Get the adjacent examination room broadcast from the neighbor room audio file, and determine the adjacent examination room broadcast adjustment parameter according to the preset wall attenuation to generate the adjusted adjacent examination room broadcast; Mix the adjusted noise and the adjusted adjacent examination room broadcast into the left and right ear clean sound signals to obtain the target left and right ear sound signals of the target seat coordinates.

2. The listening training system for English teaching according to claim 1, wherein In the target examination room, GxH reference points are arranged in a row-column grid, left and right ear sound reflection data are recorded at GxH reference points respectively, and stored as a BRIR grid library, comprising: Measuring the horizontal width of the target examination room and the longitudinal length and setting the origin of the venue coordinate system ; Based on the horizontal width , the vertical length , and the grid row and column number G and H, the horizontal interval and the vertical interval are calculated respectively; Loading preset test signal , represents the n-th sampling point signal in the preset test signal, , and playing through the examination room speaker At coordinates Left and right ear response signals are recorded at coordinates With ; left and right ear response signals with with test signal performing deconvolution operations: , , wherein, represents left ear sound reflection data, represents right ear sound reflection data, represents a discrete Fourier transform, represents an inverse discrete Fourier transform; are stored as a BRIR grid library. with organized and stored by row and column indices as a BRIR grid library.

3. The listening training system for English teaching according to claim 2, wherein Record the natural noise of the examination room at GxH reference points and save it as a noise template library; Record the broadcast content in the adjacent examination room of the target examination room and save it as a neighbor room audio file; comprising: recording a natural noise signal in the target examination room for a preset duration at GxH reference points in the target examination room , denotes the signal of the nth sampling point in the natural noise signal, 1≤n≤N; To a natural noise signal Computing the sound pressure level of the corresponding reference point As follows: , wherein, represents the number of sampling points within the preset time length; recording a broadcasting content signal for a preset duration in an adjacent examination room of the target examination room , denotes the nth sample point signal in the broadcasting content signal To a voice content signal Computing a sound pressure level As follows: , a natural noise signal and a corresponding sound pressure level is stored as a noise template library; The audio content signal and corresponding sound pressure level is stored as a field adjacent speech file.

4. The listening training system for English teaching according to claim 3, wherein Determine the three reference points corresponding to the target seat coordinates and the corresponding left and right ear sound reflection data in the BRIR grid library, and fuse to obtain the virtual reflection sound data of the target seat coordinates, comprising: determining a target seat coordinate P set by the user, ; Discretize the BRIR grid library into several triangular elements through GxH reference points; Determine the triangular element containing the target seat coordinates, and obtain the reference points A, B and C corresponding to the triangular element; The coordinates of the reference points A, B and C are determined respectively as , , ; Calculate the triangular element barycentric coordinate weight: , , , , wherein, denotes a triangle vertex barycentric coordinate weight normalization factor, , and denote the linear blending weights of the reference points A, B and C, respectively. Respectively take the left and right ear sound reflection data corresponding to reference points A, B and C, and perform linear fusion according to the triangular element barycentric coordinate weight 、 and to obtain the virtual reflection sound data of the target seat.

5. The listening training system for English teaching according to claim 4, wherein Determine the volume adjustment parameter according to the Euclidean distance between the target seat coordinates and the examination room loudspeaker, and the reference sound pressure level, comprising: According to the target seat coordinates and the pre-recorded examination room speaker coordinates ; determining the euclidean distance between the target seat coordinates and the examination hall speaker coordinates ; Determining unit distance based on prior knowledge Corresponding reference sound pressure level ; Based on the reference distance and the reference sound pressure level , the target sound pressure level is calculated according to the free field attenuation model: , wherein represents the target sound pressure level; Determine the volume adjustment parameter, as follows: , wherein represents a volume adjustment parameter.

6. The listening training system for English teaching according to claim 5, wherein Convolve the hearing training audio with the virtual reflection sound data, and apply the volume adjustment parameter to generate left and right ear clean sound signals, comprising: Acquiring hearing training audio , virtual reflected sound data of a target seat and and a volume adjustment parameter ; Perform discrete convolution operation on the virtual reflection sound data and the hearing training audio, as follows: , , wherein, denotes the amplitude of the intermediate signal after the convolution operation at the n-th sampling point, denotes the amplitude of the intermediate signal after the convolution operation at the n-th sampling point, denotes the amplitude of the hearing training audio at the n-th sampling point, denotes a time delay, denotes the amplitude of the intermediate signal after the convolution operation at the n-th sampling point, denotes the amplitude of the intermediate signal after the convolution operation at the n-th sampling point, denotes Q-1, Q being the number of sampling points in the hearing training audio, denotes a convolution operation; Will and Respectively with volume adjustment parameters Multiply them to generate clean sound signals for both ears.

7. The listening training system for English teaching according to claim 6, wherein Select the natural noise of the reference point with the smallest Euclidean distance from the target seat coordinates from the noise template library, and determine the noise adjustment parameter according to the preset target signal-to-noise ratio to generate the adjusted noise, comprising: calculating the Euclidean distance between each reference point in the noise template library and the target seat coordinates ; determining to cause a minimum reference point, and selecting a natural noise signal corresponding to the reference point , According to the left ear clean sound signal With a preset target signal-to-noise ratio Calculate a noise adjustment parameter As follows: , wherein denotes the computation of the Euclidean norm; The natural noise signal is multiplied with the noise adjustment parameter to generate the adjusted noise.

8. The listening training system for English teaching according to claim 7, wherein Get the adjacent examination room broadcast from the neighbor room audio file, and determine the adjacent examination room broadcast adjustment parameter according to the preset wall attenuation to generate the adjusted adjacent examination room broadcast, comprising: Obtaining a voice content signal from a neighboring field voice file ; Obtaining wall signal attenuation of target examination room and adjacent examination room ; wall signal attenuation amount converted into a voice adjustment parameter , ; The audio content signal is multiplied with the audio adjustment parameter to generate an adjusted adjacent test center audio.

9. The listening training system for English teaching according to claim 8, wherein Mix the adjusted noise and the adjusted adjacent examination room broadcast sound into the clean sound signals of the left and right ears to obtain target left and right ear sound signals of the target seat coordinate, including: determining a left ear clean sound signal , a right ear clean sound signal , adjusted noise , and adjusted adjacent test room announcements ; The target left ear sound signal of the target seat coordinate is calculated as follows: , wherein, a target left ear sound signal representing target seat coordinates; The target right ear sound signal of the target seat coordinate is calculated as follows: The target right ear sound signal of the target seat coordinate is calculated as follows: , wherein, a target left ear sound signal representing target seat coordinates.

Citation Information

Patent Citations

  • Visual communication advertisement design system and method

    CN120655355A

  • BR0000227A