Intelligent Sound Control Method for Multimedia Speakers Based on Big Data

By using sensor arrays in multimedia audio to collect voice and calculate user position, adjust the sound angle and volume of the audio, the problem of low volume control accuracy in the prior art is solved, and a higher degree of sound and user fit is achieved.

CN118828289BActive Publication Date: 2025-06-27SHENZHEN HUANGQING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410964470.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2025-06-27
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

The existing intelligent sound control method of multimedia audio cannot effectively compensate for the amplitude loss of sound during the propagation process, resulting in low volume control accuracy.

Method used

The intelligent sound control method of multimedia audio based on big data is adopted to collect the user's voice through the sensor array, calculate the user's orientation and straight line distance relative to the audio, and adjust the sound angle and volume of the audio based on this information to meet the user's control needs.

Benefits of technology

Through the compensation and distance calculation of multiple sensors, the accuracy of volume control is improved and the fit between the speaker and the user is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118828289B_ABST
    Figure CN118828289B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent sound control method for multimedia speakers based on big data, which relates to the field of intelligent control. The multimedia speaker has functions of adjusting the horizontal sound emission angle and the vertical sound emission angle. The intelligent sound control method includes the following steps: collecting the user's voice through a sound collection sensor array; using the phase difference of the voice collected by the sensor array to calculate the orientation of the user relative to the speaker, and determining the straight-line distance of the user relative to the speaker according to the orientation. By using the sensor array to collect sound, on the one hand, multiple sensors can compensate for missing values with each other, thus ensuring the integrity of sound data. On the other hand, multiple sensors can also determine the orientation of the user, so that the sound emission direction of the speaker can fit the orientation of the user, enhancing the usage effect. At the same time, based on the orientation, the straight-line distance can be determined, and combined with the loss function of sound propagation, the volume loss value can be compensated, increasing the accuracy of volume control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent control, and particularly to an intelligent sound control method for multimedia speakers based on big data. Background Art

[0002] With the continuous development of technology, multimedia speakers have become an indispensable part of people's lives. And with the development of voice acquisition and recognition technology, the technology of remotely controlling speakers by voice has gradually emerged.

[0003] However, in the existing intelligent sound control methods for multimedia speakers in use, due to the amplitude loss of sound during propagation, the volume size has a loss, and it is impossible to compensate for the loss based on the distance of the user, resulting in a low precision of remote voice control.

[0004] Therefore, the present invention proposes an intelligent sound control method for multimedia speakers based on big data. Summary of the Invention

[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and propose an intelligent sound control method for multimedia speakers based on big data.

[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0007] An intelligent sound control method for multimedia speakers based on big data. The multimedia speaker has functions of adjusting the horizontal sound emission angle and the vertical sound emission angle. The intelligent sound control method includes the following steps:

[0008] S1: Collect the voice of the user through a sound collection sensor array;

[0009] S2: Use the sensor array to collect the phase difference of the voice, calculate the azimuth of the user relative to the speaker, and determine the straight-line distance of the user relative to the speaker according to the azimuth;

[0010] S3: Adjust the sound emission angle of the multimedia speaker according to the azimuth of the user so that it faces the user;

[0011] S4: Process the collected voice of the user, and analyze the control requirement A of the user for the sound volume size;

[0012] S5: Adjust the volume size of the speaker according to the volume size requirement A, the sound propagation loss model in the air, and the straight-line distance of the user relative to the speaker, so that when the sound propagates to the user, the volume size is the control requirement A.

[0013] Preferably: In the step S2, the method for determining the azimuth includes the following steps:

[0014] S21A: At least set up a 3*3 sensor array, the distance between adjacent sensors in the sensor array is ΔL, and in the sensor array, arbitrarily select three horizontally arranged sensors and three vertically arranged sensors;

[0015] S22A: Establish a spatial coordinate system, and according to the positions of the sound device and the sensor array, determine that the coordinates of the three horizontally arranged sensors are respectively: The coordinates of the three vertically arranged sensors are respectively: And preset the user coordinate information as (x, y, z);

[0016] S23A: Calculate the distances between the user coordinates and the sensors respectively

[0017] S24A: For the three horizontally arranged sensors, perform permutations and combinations to obtain three combinations P1P2, P1P3, P2P3, and the sound phase differences δP1P2, δP1P3, δP2P3 collected by them are respectively. For the three horizontally arranged sensors, perform permutations and combinations to obtain three combinations P4P5, P4P6, P5P6, and the sound phase differences collected by them are respectively δP4P5, δP4P6, δP5P6;

[0018] S25A: Calculate the actual distance differences between the coordinate distance differences and the phase differences in each combination respectively, and establish a functional relationship by making an equation, L i -L o = CδP i P o , where C is the speed of sound propagation in the air; obtain six functional relationships, and then solve the functional relationships to obtain the values of x, y, z, and determine the user coordinates.

[0019] Preferably: In the S2 step, the method for determining the straight-line distance includes the following steps:

[0020] S21B: Determine the coordinates (x′, y′, z′) of the sound-emitting unit of the sound device according to the sound device structure;

[0021] S22B: Combine the user coordinates and calculate the straight-line distance according to the formula where x, y, z are the coordinates of the user.

[0022] Preferably: In the S4 step, the voice processing includes preprocessing, feature extraction, and speech recognition.

[0023] Preferably: The preprocessing includes the following steps:

[0024] S41A: Compare the sound signals collected by each sensor;

[0025] S42A: Filter the sound signal;

[0026] S43A: Traverse the sound signals of each sensor. When there is a missing signal, detect the sound signals of the remaining sensors at the missing position and fill in the missing part according to these sound signals.

[0027] Preferably, in step S4, the feature extraction includes the following steps:

[0028] S41B: Frame segmentation: The robot microphone collects voice parameters and performs frame segmentation on the voice stream. The frame length is 512 sampling points and the frame shift is 128 sampling points;

[0029] S42B: Pre-emphasis to offset the suppression of high-frequency sounds by the oral cavity. For each frame, perform the following processing: y(0) = 0.03×s(1), y(n) = s(n) - 0.97×s(n - 1), n = 1, 2,....., 51, where s(n) represents the (n + 1)-th sampling point in a frame;

[0030] S43B: Windowing: Use a Hamming window to weaken the discontinuity of the signal at the boundaries caused by frame segmentation: where n = 0, 1, ……, T - 1, T = 512.

[0031] S44B: Fast Fourier Transform: Use a radix-2 discrete Fourier transform to convert the time-domain energy to frequency-domain energy: where k = 0, 1, ……, N - 1, N = 512, and perform a flattening operation on it to obtain the energy in the real number domain: X k = |X k | 2 ;

[0032] S45B: MEL energy: Obtain the 40-dimensional MEL frequency sub-band energy through 40 MEL filter banks;

[0033] S46B: Mel log energy: Take the logarithm of each mel frequency sub-band energy: mel(i) = ln(filt(i)), i = 1, 2,......, 40;

[0034] S47B: Discrete Cosine Transform: where i = 0, 1......, M - 1; M = 40; D = 13.

[0035] Preferably, in step S5, the volume control method includes the following steps:

[0036] S51: Determine the loss function of volume loss and distance when the sound propagates, ΔD = f(ΔL), where ΔD is the volume loss and ΔL is the propagation distance;

[0037] S52: Determine the volume loss value D based on the user's distance L in combination with the loss function, D = f(L);

[0038] S53: Just adjust the volume of the audio to A + D, where A is the control requirement for the volume size.

[0039] Preferably: In the step S51, when the audio is in an open environment, the loss function is

[0040] Preferably: In the step S51, when the audio is in an indoor environment, the loss function is

[0041] The beneficial effects of the present invention are as follows:

[0042] 1. In the present invention, by using the sensor array to collect sound, on the one hand, multiple sensors can compensate for the missing values with each other, thus ensuring the integrity of the sound data. In addition, multiple sensors can also determine the orientation of the user, so that the sound emission direction of the audio can fit the orientation of the user, enhancing the use effect. At the same time, based on the orientation, the straight-line distance can also be determined, and then combined with the loss function of sound propagation, the volume loss value can be compensated, increasing the accuracy of volume control. Description of the Drawings

[0043] Figure 1 It is a flowchart of the intelligent sound control method for a multimedia audio based on big data proposed by the present invention. Detailed Embodiments

[0044] The technical solutions of the present invention will be further described in detail below in combination with the specific embodiments.

[0045] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", "connection", and "setting" should be understood in a broad sense. For example, it can be fixedly connected and set, or detachably connected and set, or integrally connected and set. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0046] Embodiment 1:

[0047] An intelligent sound control method for a multimedia audio based on big data, the multimedia audio has the functions of adjusting the horizontal sound emission angle and the vertical sound emission angle, and its intelligent sound control method includes the following steps:

[0048] S1: Collect the user's voice through the sound collection sensor array;

[0049] S2: using the sensor array to collect the phase difference of the voice, calculating the position of the user relative to the sound, and determining the straight-line distance of the user relative to the sound according to the position;

[0050] S3: adjusting the sound angle of the multimedia speaker according to the user's position so that it faces the user;

[0051] S4: Processing the collected user's voice and analyzing the user's demand A for controlling the sound volume;

[0052] S5: Adjust the volume of the speaker according to the volume demand A, the sound propagation loss model in the air, and the straight-line distance of the user relative to the speaker, so that when the sound is transmitted to the user, the volume is at the control demand A.

[0053] In the step S2, the method for determining the orientation comprises the following steps:

[0054] S21A: at least a 3*3 sensor array is provided, the distance between adjacent sensors in the sensor array is ΔL, and in the sensor array, three sensors arranged horizontally and three sensors arranged vertically are randomly selected;

[0055] S22A: Establish a spatial coordinate system, and determine the coordinates of the three transversely arranged sensors based on the positions of the sound and sensor arrays: The coordinates of the three vertically arranged sensors are: And preset the user coordinate information to (x, y, z);

[0056] S23A: Calculate the distance between the user coordinates and the sensor

[0057] S24A: For the three transversely arranged sensors, three combinations P1P2, P1P3, and P2P3 are obtained, and the sound phase differences δP1P2, δP1P3, and δP2P3 are collected respectively. For the three transversely arranged sensors, three combinations P4P5, P4P6, and P5P6 are obtained, and the sound phase differences collected are δP4P5, δP4P6, and δP5P6 respectively.

[0058] S25A: Calculate the actual distance difference of the coordinate distance difference and the phase difference in each combination, and make an equation to establish a functional relationship, L i -L o =CδP i P o , where C is the speed of sound propagation in the air; six functional relationships are obtained, and then the most functional relationship is solved to obtain the x, y, and z values ​​to determine the user's coordinates.

[0059] In step S2, the method for determining the straight-line distance comprises the following steps:

[0060] S21B: Determine the coordinates (x′, y′, z′) of the sound unit of the sound according to the sound structure;

[0061] S22B: Combine the user's coordinates and calculate the straight-line distance according to the formula Where x, y, and z are the coordinates of the user.

[0062] Embodiment 2:

[0063] In the step S2, the method for determining the orientation comprises the following steps:

[0064] S21A: at least a 3*3 sensor array is provided, the distance between adjacent sensors in the sensor array is ΔL, and in the sensor array, three sensors arranged horizontally and three sensors arranged vertically are randomly selected;

[0065] S22A: Establish a spatial coordinate system, and determine the coordinates of the three transversely arranged sensors based on the positions of the sound and sensor arrays: The coordinates of the three vertically arranged sensors are: And preset the user coordinate information to (x, y, z);

[0066] S23A: Calculate the distance between the user coordinates and the sensor

[0067] S24A: For the three transversely arranged sensors, three combinations P1P2, P1P3, and P2P3 are obtained, and the sound phase differences δP1P2, δP1P3, and δP2P3 are collected respectively. For the three transversely arranged sensors, three combinations P4P5, P4P6, and P5P6 are obtained, and the sound phase differences collected are δP4P5, δP4P6, and δP5P6 respectively.

[0068] S25A: Calculate the actual distance difference of the coordinate distance difference and the phase difference in each combination, and make an equation to establish a functional relationship, L i -L o =CδP i P o , where C is the speed of sound propagation in the air; six functional relationships are obtained, and then the most functional relationship is solved to obtain the x, y, and z values ​​to determine the user's coordinates.

[0069] In step S2, the method for determining the straight-line distance comprises the following steps:

[0070] S21B: Determine the coordinates (x′, y′, z′) of the sound unit of the sound according to the sound structure;

[0071] S22B: Combine the user coordinates and calculate the straight-line distance according to the formula where x, y, and z are the coordinates of the user.

[0072] In the step S4, the voice processing includes preprocessing, feature extraction, and speech recognition.

[0073] The preprocessing includes the following steps:

[0074] S41A: Compare the sound signals collected by each sensor;

[0075] S42A: Filter the sound signals;

[0076] S43A: Traverse the sound signals of each sensor. When there is a missing signal, detect the sound signals of the remaining sensors at the missing position and fill the missing part according to these sound signals.

[0077] In the step S4, the feature extraction includes the following steps:

[0078] S41B: Frame segmentation: The robot microphone collects voice parameters and performs frame segmentation on the voice stream. The frame length is 512 sampling points, and the frame shift is 128 sampling points;

[0079] S42B: Pre-emphasis to offset the suppression of high-frequency sounds by the oral cavity. For each frame, perform the following processing: y(0) = 0.03 × s(1), y(n) = s(n) - 0.97 × s(n - 1), n = 1, 2,......, 51, where s(n) represents the (n + 1)-th sampling point in a frame;

[0080] S43B: Windowing: Use a Hamming window to reduce the discontinuity of the signal at the boundaries caused by frame segmentation: where n = 0, 1, ……, T - 1, T = 512.

[0081] S44B: Fast Fourier transform: Use the radix-2 discrete Fourier transform to convert the time-domain energy to frequency-domain energy: where k = 0, 1, ……, N - 1, N = 512, and perform a flattening operation on it to obtain the energy in the real number domain: X k = |X k | 2 ;

[0082] S45B: MEL energy: Obtain the 40-dimensional MEL frequency sub-band energy through 40 MEL filter banks;

[0083] S46B: Mel Logarithmic Energy: Take the logarithm of the energy of each mel frequency sub-band: mel(i) = ln(filt(i)), where i = 1, 2,......, 40;

[0084] S47B: Discrete Cosine Transform: where i = 0, 1......, M - 1; M = 40; D = 13.

[0085] Example 3:

[0086] In the step S2, the method for determining the orientation includes the following steps:

[0087] S21A: Set at least a 3*3 sensor array. The distance between adjacent sensors in the sensor array is ΔL. In the sensor array, arbitrarily select three horizontally arranged sensors and three vertically arranged sensors;

[0088] S22A: Establish a spatial coordinate system, and according to the positions of the sound device and the sensor array, determine that the coordinates of the three horizontally arranged sensors are respectively: The coordinates of the three vertically arranged sensors are respectively: And preset the user coordinate information as (x, y, z);

[0089] S23A: Calculate the distances between the user coordinates and the sensors respectively

[0090] S24A: Perform permutations and combinations on the three horizontally arranged sensors to obtain three combinations P1P2, P1P3, P2P3. The sound phase differences δP1P2, δP1P3, δP2P3 collected by them are respectively. Perform permutations and combinations on the three horizontally arranged sensors to obtain three combinations P4P5, P4P6, P5P6. The sound phase differences collected by them are respectively δP4P5, δP4P6, δP5P6;

[0091] S25A: Calculate the actual distance differences between the coordinate distance differences and the phase differences in each combination respectively, and establish a functional relationship by making an equation, L i -L o = CδP i P o , where C is the speed of sound propagation in the air; obtain six functional relationships, and then solve the functional relationships to obtain the values of x, y, z and determine the user coordinates.

[0092] In the step S2, the method for determining the straight-line distance includes the following steps:

[0093] S21B: Determine the coordinates (x′, y′, z′) of the sound-emitting unit of the sound device according to the sound device structure;

[0094] S22B: Calculate the straight-line distance according to the formula in combination with the user's coordinates Where x, y, and z are the coordinates of the user.

[0095] In the step S4, the voice processing includes preprocessing, feature extraction, and speech recognition.

[0096] The preprocessing includes the following steps:

[0097] S41A: Compare the sound signals collected by each sensor;

[0098] S42A: Filter the sound signals;

[0099] S43A: Traverse the sound signals of each sensor. When there is a missing value, detect the sound signals of the remaining sensors at the missing position, and fill the missing value according to the sound signal.

[0100] In the step S4, the feature extraction includes the following steps:

[0101] S41B: Frame division: The robot microphone collects voice parameters and performs frame division on the voice stream. The frame length is 512 sampling points, and the frame shift is 128 sampling points;

[0102] S42B: Pre-emphasis to offset the suppression of high-frequency sounds by the oral cavity. For each frame, perform the following processing: y(0) = 0.03 × s(1), y(n) = s(n) - 0.97 × s(n - 1), n = 1, 2,......, 51, where s(n) represents the (n + 1)-th sampling point in a frame;

[0103] S43B: Windowing: Use a Hamming window to weaken the discontinuity of the signal at the boundary caused by frame division: Where n = 0, 1, ……, T - 1, T = 512.

[0104] S44B: Fast Fourier transform: Use the radix-2 discrete Fourier transform to convert the time-domain energy into frequency-domain energy: Where k = 0, 1, ……, N - 1, N = 512, and perform a flattening operation on it to obtain the energy in the real number domain: X k = |X k | 2 ;

[0105] S45B: MEL energy: Obtain the 40-dimensional MEL frequency sub-band energy through 40 MEL filter banks;

[0106] S46B: Mel logarithmic energy: for each mel frequency subband energy, de-contrast: mel(i) = ln(filt(i)), i = 1, 2, ..., 40;

[0107] S47B: Discrete Cosine Transform: Among them, i=0, 1..., M-1; M=40; D=13.

[0108] In step S5, the volume control method comprises the following steps:

[0109] S51: Determine the loss function of volume loss and distance during sound propagation ΔD=f(ΔL), where ΔD is the volume loss and ΔL is the propagation distance;

[0110] S52: Determine a volume loss value D according to the user distance L and the loss function, where D=f(L);

[0111] S53: Adjust the audio volume to A+D, where A is the volume control requirement.

[0112] In the step S51, the loss function is:

[0113] Embodiment 4:

[0114] In the step S2, the method for determining the orientation comprises the following steps:

[0115] S21A: at least a 3*3 sensor array is provided, the distance between adjacent sensors in the sensor array is ΔL, and in the sensor array, three sensors arranged horizontally and three sensors arranged vertically are randomly selected;

[0116] S22A: Establish a spatial coordinate system, and determine the coordinates of the three transversely arranged sensors based on the positions of the sound and sensor arrays: The coordinates of the three vertically arranged sensors are: And preset the user coordinate information to (x, y, z);

[0117] S23A: Calculate the distance between the user coordinates and the sensor

[0118] S24A: For the three transversely arranged sensors, three combinations P1P2, P1P3, and P2P3 are obtained, and the sound phase differences δP1P2, δP1P3, and δP2P3 are collected respectively. For the three transversely arranged sensors, three combinations P4P5, P4P6, and P5P6 are obtained, and the sound phase differences collected are δP4P5, δP4P6, and δP5P6 respectively.

[0119] S25A: Calculate the actual distance differences of the coordinate distance differences and phase differences in each combination respectively, and establish a functional relationship by making an equation with them, L i -L o = CδP i P o , where C is the speed of sound propagation in air; obtain six functional relationships, and then solve the functional relationships to obtain the values of x, y, and z to determine the coordinates of the user.

[0120] In the step S2, the method for determining the straight-line distance includes the following steps:

[0121] S21B: Determine the coordinates (x′, y′, z′) of the sound-emitting unit of the speaker according to the speaker structure;

[0122] S22B: Combine the user coordinates and calculate the straight-line distance according to the formula where x, y, and z are the coordinates of the user.

[0123] In the step S4, the speech processing includes preprocessing, feature extraction, and speech recognition.

[0124] The preprocessing includes the following steps:

[0125] S41A: Compare the sound signals collected by each sensor;

[0126] S42A: Perform filtering processing on the sound signals;

[0127] S43A: Traverse the sound signals of each sensor. When there is a missing signal, detect the sound signals of the remaining sensors at the missing position, and fill the missing signal according to this sound signal.

[0128] In the step S4, the feature extraction includes the following steps:

[0129] S41B: Frame division: The robot microphone collects speech parameters and performs frame division processing on the speech stream. The frame length is 512 sampling points, and the frame shift is 128 sampling points;

[0130] S42B: Pre-emphasis to offset the suppression effect of the oral cavity on the high-frequency part of the sound. For each frame, perform the following processing: y(0) = 0.03×s(1), y(n) = s(n) - 0.97×s(n - 1), n = 1, 2,......, 51, where s(n) represents the (n + 1)-th sampling point in a frame;

[0131] S43B: Windowing: Use a Hamming window to weaken the discontinuity of the signal at the boundary caused by frame division: where n = 0, 1, ……, T - 1, T = 512.

[0132] S44B: Fast Fourier Transform: Convert the time-domain energy to frequency-domain energy using the radix-2 discrete Fourier transform: where k = 0, 1, ……, N - 1, N = 512, and perform a flattening operation on it to obtain the energy in the real number domain: X k = |X k | 2 ;

[0133] S45B: MEL Energy: Obtain the 40-dimensional MEL frequency sub-band energy through 40 MEL filter banks;

[0134] S46B: Mel Logarithmic Energy: Take the logarithm of each mel frequency sub-band energy: mel(i) = ln(filt(i)), i = 1, 2,......, 40;

[0135] S47B: Discrete Cosine Transform: where i = 0, 1......, M - 1; M = 40; D = 13.

[0136] In the step S5, the volume control method includes the following steps:

[0137] S51: Determine the loss function of volume loss and distance ΔD = f(ΔL) when the sound propagates, where ΔD is the volume loss and ΔL is the propagation distance;

[0138] S52: Determine the volume loss value D according to the user's distance L and the loss function, D = f(L);

[0139] S53: Adjust the volume of the speaker to A + D, where A is the control requirement for the volume size.

[0140] In the step S51, the loss function is

[0141] As mentioned above, the above is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A multimedia audio intelligent sound control method based on big data, characterized in that: The multimedia speaker has the functions of adjusting the horizontal sound angle and the vertical sound angle, and the intelligent sound control method thereof comprises the following steps: S1: Collect the user's voice through the sound collection sensor array; S2: using the sensor array to collect the phase difference of the voice, calculating the position of the user relative to the sound, and determining the straight-line distance of the user relative to the sound according to the position; S3: adjusting the sound angle of the multimedia speaker according to the user's position so that it faces the user; S4: Processing the collected user's voice and analyzing the user's demand A for controlling the sound volume; S5: adjusting the volume of the sound system according to the volume demand A, the sound propagation loss model in the air, and the straight-line distance of the user relative to the sound system, so that when the sound is transmitted to the user, the volume is at the control demand A; In the step S2, the method for determining the orientation comprises the following steps: S21A: at least a 3*3 sensor array is provided, the distance between adjacent sensors in the sensor array is ΔL, and in the sensor array, three sensors arranged horizontally and three sensors arranged vertically are randomly selected; S22A: Establish a spatial coordinate system, and determine the coordinates of the three transversely arranged sensors based on the positions of the sound and sensor arrays: The coordinates of the three vertically arranged sensors are: And preset the user coordinate information to (x, y, z); S23A: Calculate the distance between the user coordinates and the sensor S24A: For the three transversely arranged sensors, three combinations P1P2, P1P3, and P2P3 are obtained, and the sound phase differences δP1P2, δP1P3, and δP2P3 are collected respectively. For the three transversely arranged sensors, three combinations P4P5, P4P6, and P5P6 are obtained, and the sound phase differences collected are δP4P5, δP4P6, and δP5P6 respectively. S25A: Calculate the actual distance difference of the coordinate distance difference and the phase difference in each combination, and make an equation to establish a functional relationship, L i -L o =CδP i P o , where C is the speed of sound propagation in the air; get six functional relationships, then solve the most functional relationship to get the x, y, z values ​​and determine the user's coordinates; Among them, δ is the physical quantity representing the phase difference; L0 and L1 represent the distances between different sensors and the user, and are based on the distance values ​​calculated in step S23A; P i and P0 is the sensor identification.

2. The intelligent sound control method for multimedia audio based on big data according to claim 1 is characterized in that: In step S2, the method for determining the straight-line distance comprises the following steps: S21B: Determine the coordinates (x′, y′, z′) of the sound unit of the sound according to the sound structure; S22B: Combine the user's coordinates and calculate the straight-line distance according to the formula Where x, y, and z are the coordinates of the user.

3. The intelligent sound control method for multimedia audio based on big data according to claim 1 is characterized in that: In the step S4, speech processing includes preprocessing, feature extraction and speech recognition.

4. The intelligent sound control method for multimedia audio based on big data according to claim 3 is characterized in that: The pre-processing comprises the following steps: S41A: Compare the sound signals collected by each sensor; S42A: filtering the sound signal; S43A: Traverse the sound signals of each sensor, and when there is a missing part, detect the sound signals of the remaining sensors at the missing part, and fill the missing part according to the sound signals.

5. The intelligent sound control method for multimedia audio based on big data according to claim 4 is characterized in that: In the step S4, feature extraction includes the following steps: S41B: Framing: The robot microphone collects voice parameters and performs framing on the voice stream. The frame length is 512 sampling points and the frame shift is 128 sampling points. S42B: pre-emphasis to offset the suppression of the high-frequency sound by the oral cavity. For each frame, the following processing is performed: y(0)=0.03×s(1), y(n)=s(n)-0.97×s(n-1), n=1, 2, ..., 51, where s(n) represents the n+1th sampling point in a frame; S43B: Windowing: Use a Hamming window to reduce the discontinuity of the signal at the boundary caused by framing: Where n = 0, 1, ..., T-1, T = 512. S44B: Fast Fourier Transform: Convert time domain energy to frequency domain energy using a base-2 discrete Fourier transform: Where k = 0, 1, ..., N-1, N = 512, and the energy in the real number domain is obtained by flattening it: X k =|X k | 2 ; S45B: MEL energy: Through 40 MEL filter banks, 40-dimensional MELMEL frequency subband energy is obtained; S46B: Mel logarithmic energy: for each mel frequency subband energy, de-interpolate: mel(i) = ln(filt(i)), i = 1, 2, ..., 40; S47B: Discrete Cosine Transform: Among them, i=0, 1,..., M-1; M=40; D=13.

6. The intelligent sound control method for multimedia audio based on big data according to claim 1 is characterized in that: In step S5, the volume control method comprises the following steps: S51: Determine the loss function of volume loss and distance during sound propagation ΔD=f(ΔL), where ΔD is the volume loss and ΔL is the propagation distance; S52: Determine a volume loss value D according to the user distance L and the loss function, where D=f(L); S53: Adjust the audio volume to A+D, where A is the volume control requirement.

7. The intelligent sound control method for multimedia audio based on big data according to claim 6 is characterized in that: In the step S51, when the sound is in an open environment, the loss function is 8. The intelligent sound control method for multimedia audio based on big data according to claim 7 is characterized in that: In the step S51, when the sound system is in an indoor environment, the loss function is

Citation Information

Patent Citations

  • Interactive multifunctional elderly care robot recognition system

    CN111862950A

  • Operation control system based on intelligent man-machine interaction

    CN112017658A