Adaptive sound field method, apparatus and device, and computer readable storage medium
By calculating sound wave paths and phase modulation, combined with emotion analysis, the problem of fixed listening positions in traditional audio systems has been solved, enabling intelligent speakers to achieve adaptive sound fields, improving sound quality and hearing protection, and adapting to user needs in different scenarios.
Patent Information
- Application Number
- CN202511019468.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional audio systems suffer from problems such as fixed optimal listening positions, poor adaptability to sound field environments, low levels of intelligence, and insufficient hearing protection mechanisms, resulting in a poor user experience.
By calculating the direct and reflection paths of sound waves to the user's ear, the phase and gain of the sound waves are adjusted using a reverse calculation method to form a target sound field, and then combined with an emotion analysis model for intelligent speaker control.
It achieves an immersive experience with optimal sound quality in any location, enhances users' spatial awareness and hearing protection, improves intelligence, and adapts to user needs in different scenarios.
Smart Images

Figure CN120802625A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of data processing, and in particular to an adaptive sound field method, device, equipment and computer readable storage medium. BACKGROUND
[0002] Current sound products have the following characteristics: Fixed optimal listening position: traditional high-fidelity sound (Hi-Fi) systems require complex sound calibration, and the optimal listening position is usually only one, which cannot provide the best experience for all people in the room.
[0003] Poor sound field environment adaptability: the placement of the sound and the change of furniture in the room will seriously affect the sound quality, and each adjustment needs to be recalibrated, which is very inconvenient.
[0004] “Pseudo-intelligent” of smart sound: current smart sound can be controlled by voice, but it has no information about “who” is giving the order, how many people are in the room, and how the state of everyone is, and it is only a passive order receiver, which cannot actively optimize and adapt to the scene.
[0005] Ignoring hearing health: users, especially young users, tend to listen to music or watch movies for a long time at high volume, and current devices lack effective and intelligent hearing protection mechanisms.
[0006] That is, in the prior art, the optimal listening position is usually only one, which cannot provide the best experience for all people in the room; the placement of the sound and the change of furniture in the room will seriously affect the sound quality, and each adjustment needs to be recalibrated, which is very inconvenient; the sound system is less intelligent; and the hearing protection mechanism is poor. SUMMARY
[0007] According to embodiments of the present application, an adaptive sound field scheme is provided, which can calculate the direct path of sound waves reaching the user's ear and the path after reflection via different walls, and cooperates with the algorithm of the present disclosure to control the sound waves on the path, so that the sound waves on all paths reach the listener's ear at the same time and with the same phase. It greatly enhances the sound pressure and clarity of the target position, can create a kind of completely surrounded by sound, and has a strong sense of space immersive experience, greatly improving the user experience.
[0008] In a first aspect of the present application, an adaptive sound field method is provided. The method comprises:
[0009] obtaining target data;
[0010] determining a target sound wave path according to the target data;
[0011] According to the target sound wave path, a target playing operation is performed to form a target sound field.
[0012] Further, the target data includes:
[0013] target physical environment feature data, target acoustic feature data, and / or target user feature data.
[0014] Further, the determining the target sound wave path according to the target data includes:
[0015] According to the target data, all sound wave paths from each loudspeaker to the user position are calculated respectively to obtain a sound wave path set;
[0016] Based on the sound wave path set and the number of loudspeakers, a driving signal vector is calculated in a reverse calculation manner.
[0017] Based on the driving signal vector, the target sound wave path is determined.
[0018] Further, the calculating all sound wave paths from each loudspeaker to the user position according to the target data includes:
[0019] According to the target data, all sound wave paths from each loudspeaker to the user position are calculated respectively by ray tracing method; wherein the sound wave path includes reflection path, diffraction path, and / or absorption path.
[0020] Further, the calculating the driving signal vector based on the sound wave path set and the number of loudspeakers in a reverse calculation manner includes:
[0021] Based on the sound wave path set and the number of loudspeakers, the driving signal vector is calculated in a reverse calculation manner through the following formula:
[0022] X = (H^H H + λI)^-1 H^H P_target
[0023] Wherein, H is a transfer matrix, which is set according to the number of sound wave paths and the number of loudspeakers;
[0024] λ is a regularization coefficient;
[0025] ^H is a conjugate transpose;
[0026] P_target is the user target sound pressure in the target user feature data;
[0027] I is an identity matrix.
[0028] Further, the performing the target playing operation to form the target sound field according to the target sound wave path includes:
[0029] inputting the target user feature data into a trained sentiment analysis model to obtain a sentiment category of the current user;
[0030] According to the sentiment category and a preset hierarchical intervention strategy, a target playing operation is performed through the target sound wave path to form a target sound field.
[0031] Further, the sentiment analysis model can be trained in the following manner:
[0032] A training sample set is generated, wherein the training sample includes a script file with labeled information; the labeled information is a sentiment label;
[0033] The sentiment analysis model is trained using samples in the training sample set, taking the script file as input and the sentiment label as output, and when the uniformity rate of the output sentiment label and the labeled sentiment label meets a preset threshold, the training of the sentiment analysis model is completed.
[0034] In a second aspect of the present application, an adaptive sound field device is provided. The device comprises:
[0035] An acquisition module is configured to acquire target data;
[0036] A determination module is configured to determine a target sound wave path according to the target data;
[0037] An execution module is configured to perform a target playing operation to form a target sound field according to the target sound wave path.
[0038] In a third aspect of the present application, an electronic device is provided. The electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.
[0039] In a fourth aspect of the present application, a computer readable storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the method according to the first aspect of the present application.
[0040] The adaptive sound field method provided by the embodiments of the present application greatly enhances the sound pressure and clarity of the target position, can create a kind of completely surrounded by sound, very spatial sense immersive experience, and greatly improves the user experience.
[0041] It is to be understood that the description in the summary is not intended to identify key or essential features of embodiments of the application, nor is it intended to limit the scope of the application. Other features, aspects, and advantages of the application will become apparent from the following description, when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0042] The above and other features, aspects, and advantages of the present embodiments will become more apparent with reference to the following detailed description when considered in conjunction with the accompanying drawings. In the drawings, the same or like reference numbers can indicate the same or like elements, in which:
[0043] Figure 1 a flowchart of an adaptive sound field method according to embodiments of the present application;
[0044] Figure 2 a block diagram of an adaptive sound field device according to embodiments of the present application;
[0045] Figure 3 a structural schematic diagram of a terminal device or a server suitable for implementing embodiments of the present application. DETAILED DESCRIPTION
[0046] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some, but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0047] In addition, the term "and / or" in this document is only to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this document generally represents an "or" relationship between the front and rear associated objects.
[0048] Figure 1 A flowchart of an adaptive sound field method according to embodiments of the present disclosure is shown. The method comprises:
[0049] S110, obtaining target data.
[0050] The target data includes target physical environment feature data, target acoustic feature data, and / or target user feature data.
[0051] In some embodiments, the user environment can be scanned by means of ultra-wideband and / or laser radar scanning, a room model is established, obstacles are marked, and target physical environment feature data is obtained.
[0052] The target acoustic feature data can be obtained by recognizing the current user demand and environmental sound through a microphone array;
[0053] The target user feature data such as heart rate (HR), heart rate variability (HRV), galvanic skin response (GSR), and / or respiratory rate can be obtained by connecting the user's smart watch / bracelet through Bluetooth and / or Wi-Fi.
[0054] S120, determining a target sound wave path according to the target data.
[0055] The reflection of the traditional sound wall and ceiling is interference. In the present disclosure, it can be regarded as a resource that can be utilized. By calculating the direct path of the sound wave reaching the user's ear and the path after reflection via different wall surfaces, the sound waves on the path are cooperatively regulated so as to reach the listener's ear at the same time and with the same phase.
[0056] In some embodiments, the path calculation of the sound wave can be performed in the manner of ray tracing:
[0057] A plurality of rays (such as every 10° angle interval) are emitted from each loudspeaker unit, and the direct path and the reflected path (the number of reflections is ≤3 times to balance the accuracy and the amount of calculation) are tracked;
[0058] Reflection path calculation: according to the wall surface normal vector and the incident angle, the reflection angle is calculated according to the law of specular reflection.
[0059] Path length calculation:
[0060] Direct path length: D_direct = √[(x_u-x_s)^2+(y_u-y_s)^2+(z_u-z_s)^2];
[0061] Reflection path length: D_reflect = D1+D2;
[0062] Wherein, D1 is the distance from the loudspeaker to the reflection point;
[0063] D2 is the distance from the reflection point to the user;
[0064] (x_u, y_u, z_u) is the coordinate of the user (User) in the three-dimensional space;
[0065] x_u is the position of the user on the X-axis in the room coordinate system;
[0066] y_u is the position of the user on the Y-axis in the room coordinate system;
[0067] z_u is the position of the user on the Z-axis in the room coordinate system (usually representing the height);
[0068] (x_s, y_s, z_s) is the coordinate of the speaker (Source / Speaker) in three-dimensional space;
[0069] x_s is the position of the speaker on the X-axis in the room coordinate system;
[0070] y_s is the position of the speaker on the Y-axis in the room coordinate system;
[0071] z_s is the position of the speaker on the Z-axis in the room coordinate system.
[0072] Repeat the above operation to calculate all sound wave paths from each speaker to the user position, respectively, to obtain a sound wave path set;
[0073] It should be noted that a direct sound path can include N one-time, two-time or even multiple reflection sound paths.
[0074] In some embodiments, based on the sound wave path set and the number of speakers, the driving signal vector is calculated in a reverse calculation manner.
[0075] Specifically, the driving signal vector X (size M x 1) is solved to satisfy H·X=P_target (least square solution):
[0076] X=(H^H H+λI)^{-1}H^H P_target
[0077] Wherein, H is a transfer matrix, which is set according to the number of sound wave paths and the number of speakers;
[0078] λ is a regularization coefficient;
[0079] ^H is the conjugate transpose;
[0080] P_target is the user target sound pressure in the target user feature data;
[0081] I is the unit matrix.
[0082] Further, the driving signal vector is controlled in phase and gain to cooperatively control the sound waves on all paths so that they reach the listener's ear at the same time with the same phase:
[0083] The complex form of the driving signal X is:
[0084] a_m·e^{jφ_m}
[0085] Directly give the gain a_m and phase offset φ_m of the mth speaker.
[0086] Specifically, the environment can include M speakers, and X is a vector containing M elements:
[0087] X = [x1, x2,..., x_m,..., x_M]
[0088] Each element x_m in X, which is a complex number itself, includes a real part and an imaginary part, and can be written in Cartesian form:
[0089] x_m = Re(x_m) + j * Im(x_m)
[0090] where j is the imaginary unit (√-1);
[0091] Further, the complex number x_m is analyzed to obtain the gain a_m and the phase For any complex number, in addition to the Cartesian form of "real part + imaginary part", it can also be expressed in polar form. The two forms are completely equivalent:
[0092]
[0093] where a_m and are the required gain and phase. They can be calculated from the real and imaginary parts of x_m by the following standard formulas:
[0094] Calculate the gain a_m (i.e. the "modulus" of the complex number)
[0095] The gain a_m is the strength or amplitude of the signal, equal to the modulus of the complex number x_m;
[0096] a_m = |x_m| = √[(Re(x_m))^2 + (Im(x_m))^2]
[0097] Calculate the phase shift φ_m (i.e. the "argument" of the complex number)
[0098] The phase shift φ_m is the advance or delay of the signal in time, equal to the argument of the complex number x_m.
[0099] The calculation formula is:
[0100] φ_m = arctan(Im(x_m) / Re(x_m))
[0101] It should be noted that in practical applications, the atan2(Im(x_m), Re(x_m)) function is usually used for calculation to ensure that the angle falls in the correct quadrant. Further, the phase synchronization is implemented:
[0102] By adjusting φ_m, the phase difference of all path sound waves at the user position approaches an integer multiple of 2π.
[0103] The following is an example:
[0104] Suppose there is a scene containing multiple speakers. Take the mth speaker in the scene as an example:
[0105] After the first step of complex matrix calculation, the element x_m of the driving vector X corresponding to the mth speaker is obtained:
[0106] x_m = 3 + 4j
[0107] That is, the real part Re(x_m) = 3; the imaginary part Im(x_m) = 4
[0108] At this time, according to the formula of the second step, it is analyzed:
[0109] Calculate the gain a_m:
[0110] a_m = √(3 2 +4 2 ) = √(9+16) = √25 = 5
[0111] Calculate the phase :
[0112] φ_m = arctan(4 / 3) ≈ 53.13° (or 0.927 radians)
[0113] According to the above calculation results:
[0114] Issue an instruction to the mth speaker: set the gain (amplitude) of the driving signal to 5 units, and let its phase advance by 53.13 degrees.
[0115] At this time, the same operation is performed on each complex element in the driving vector X, so that the gain and phase of each speaker that needs to be accurately adjusted are obtained.
[0116] In some embodiments, by precisely controlling the emission time difference of the speaker units, all path sound waves arrive at the user position at the same time:
[0117] Time difference calculation:
[0118] For the mth speaker, the emission time offset is:
[0119] Δt_m = T_direct_min - T_{m,direct}
[0120] Wherein, T_direct_min is the minimum delay in all direct paths of the loudspeakers (the time of flight required for the sound to travel in a straight line from the mth loudspeaker to the user's ear);
[0121] T is Time;
[0122] m is the number of the mth loudspeaker unit;
[0123] Direct is Direct Path;
[0124] If there are currently M loudspeakers, then at any time, a set of M time values T_{1,direct}, T_{2,direct},..., T_{M,direct} will be calculated.
[0125] Further, T_direct_min can be calculated as follows:
[0126] T_{m,direct} = D_{m,direct} / v_sound
[0127] Wherein, v_sound is a physical constant, about 343 meters per second, which will change slightly with temperature and humidity, and can be compensated according to the actual application scenario;
[0128] D_{m,direct} is the direct distance, that is, the straight-line distance from the mth loudspeaker to the user; it can be calculated by the following formula:
[0129] D_direct = √[(x_u - x_s)^2 +...]
[0130] Specifically, the direct times T_{1,direct}, T_{2,direct},... of all loudspeakers to the user are calculated, and the minimum value T_direct_min is found from them. This minimum value corresponds to the loudspeaker closest to the user.
[0131] Calculate the time difference:
[0132] The purpose of the formula Δt_m = T_direct_min - T_{m,direct} is to calculate how much slower each loudspeaker's direct time is relative to the "fastest" one.
[0133] For the closest loudspeaker, T_{m,direct} is equal to T_direct_min:
[0134] Δt_m = 0
[0135] For all the rest, T_{m,direct} is greater than T_direct_min, so the calculated Δt_m will be negative.
[0136] Perform compensation: Δt_m is actually a "relative emission time offset". This offset needs to be translated into an actual digital delay.
[0137] The farthest speaker (with the largest T_{m,direct}) gets the largest negative offset in absolute value. Take it as a reference, make it sound first (delay 0), and all the speakers closer to it need to be applied a positive digital delay, i.e. "wait a bit" before sounding.
[0138] The amount of delay to be applied is exactly equal to the absolute value of Δt_m (relative to the farthest speaker), and the closest speaker needs to wait the longest.
[0139] Let's illustrate with an example:
[0140] Suppose a scene with 3 speakers, with the following direct times calculated:
[0141] T_{1,direct} = 3 ms
[0142] T_{2,direct} = 5 ms
[0143] T_{3,direct} = 8 ms <- farthest, find T_direct_min = 3 ms (from speaker 1).
[0144] Calculate Δt_m for each speaker:
[0145] Δt1 = 3 ms - 3 ms = 0 ms
[0146] Δt2 = 3 ms - 5 ms = -2 ms
[0147] Δt3 = 3 ms - 8 ms = -5 ms
[0148] The DSP (Digital Signal Processor) performs delay compensation, taking the farthest speaker 3 as a reference (delay 0):
[0149] Speaker 3: needs to sound 5 ms in advance, delay 0 ms.
[0150] Speaker 2: needs to sound 2 ms in advance, delay 3 ms (5 ms - 2 ms).
[0151] Speaker 1: does not need to sound in advance, so the delay is 5 ms (5 ms - 0 ms).
[0152] Final effect:
[0153] Speaker 3 emits sound at t=0ms, travels for 8ms, and arrives at t=8ms.
[0154] Speaker 2 emits sound at t=3ms, travels for 5ms, and arrives at t=8ms.
[0155] Speaker 1 emits sound at t=5ms, travels for 3ms, and arrives at t=8ms.
[0156] In summary, the direct sound waves of all speakers arrive at the user's ear at the same moment at t=8ms, achieving perfect time synchronization. Hardware implementation:
[0157] Apply a digital delay line (precision ≤10μs) to each speaker channel through a DSP to compensate for Δt_m.
[0158] Combine the phase offset φ_m of the driving signal to achieve dual-phase calibration in the time domain and the frequency domain.
[0159] Based on the above method, by accurately controlling the time (phase) and intensity (amplitude) of sound emission of each unit, the sound waves form "constructive interference" at the position of the target user's ear and "destructive interference" in other non-target areas; wave peaks meet wave peaks, and the amplitude is enhanced (constructive interference); wave peaks meet wave troughs, and the amplitude is reduced (destructive interference).
[0160] Further, it also includes:
[0161] When the user moves (UWB positioning accuracy ±10cm), recalculate the path and update the H matrix every 100ms. Use the iterative least squares method (such as the LMS algorithm) to reuse the previous solution as the initial value to reduce the calculation delay.
[0162] S130, according to the target sound wave path, performing a target playing operation to form a target sound field.
[0163] The sensitivity and fatigue of the human ear to different frequencies are different (as shown in the famous equal-loudness contour), and long-term exposure to high or medium-high frequencies is more likely to cause hearing damage than exposure to low frequencies.
[0164] In the present disclosure, the playing of sound waves can be intelligently regulated in the following way:
[0165] Obtain the volume (SPL), spectral analysis results, and played duration of the current playing content. Construct a "dose accumulator" according to the volume (SPL), spectral analysis results, and played duration; the dose accumulator is used to describe the current state of the user, for example, when the accumulated value approaches the preset safety threshold (for example, 75% of the daily safety dose), the system is ready to start intervention (adjustment strategy);
[0166] Input the target user feature data into the trained emotion analysis model to obtain the emotion category of the current user;
[0167] According to the emotion category and the preset hierarchical intervention strategy, execute the target playing operation through the target sound wave path to form a target sound field;
[0168] The emotion analysis model can be trained in the following way:
[0169] Generate a training sample set, wherein the training sample includes a script file with labeled information; the labeled information is an emotion label;
[0170] Use the samples in the training sample set to train the emotion analysis model, taking the script file as input and the emotion label as output; when the unified rate of the output emotion label and the labeled emotion label meets the preset threshold, the training of the emotion analysis model is completed;
[0171] For example, [high HRV + low respiratory rate + night] -> "state: relaxation / sleep aid";
[0172] [high HR + low HRV + rapid breathing] -> "state: excited / stress / exercise";
[0173] [microphone detects multiple people talking] -> "scene: social gathering";
[0174] [microphone detects baby crying] -> "scene: needs to be soothed";
[0175] Further, the hierarchical intervention strategy includes:
[0176] First-level intervention (fine tuning): slightly adjust the dynamic equalizer (Dynamic EQ). Without affecting the core melody and vocals, gradually attenuate the frequency band that is most likely to cause auditory fatigue (usually 2kHz-6kHz) at an amplitude that is not easily perceived by the human ear (for example, -0.5dB);
[0177] Second-level intervention (dynamic compression): if the dose continues to rise, a slight dynamic range compression will be applied. That is, slightly raise the volume of the quieter part while compressing the peak part of the loudest part. In order to maintain the overall average loudness feeling without reducing it while reducing the maximum sound pressure level;
[0178] Third-level intervention (total gain smooth down-regulation): When the above measures still cannot control the dose, the overall volume is smoothly reduced at an extremely slow speed (for example, 0.2 dB per minute).
[0179] Fourth-level intervention (active notification): When the daily dose reaches 100%, a reminder will be sent through voice or App push: "You have enough music time today, please give your ears a rest, or switch to a more relaxing mode."
[0180] According to the embodiments of the present disclosure, the following technical effects are achieved:
[0181] By calculating the direct path of sound waves reaching the user's ear and the path after reflection via different walls, the sound waves on the path are regulated by the algorithm of the present disclosure, so that the sound waves on all paths reach the listener's ear at the same time and with the same phase. The sound pressure and clarity of the target position are greatly enhanced, creating a fully immersive experience with a strong sense of space, greatly improving the user experience.
[0182] At the same time, through the playing method of the present disclosure, it is no longer dependent on explicit voice instructions ("play sad music") of the user, and the potential needs of the user can be inferred by comprehensive analysis of various unstructured data. For example, physiological indicators (heart rate, respiration) are strong correlation signals of emotions, and environmental sound defines the context of the scene. The algorithm combines the clues to form a comprehensive judgment of the current "scene" and matches the most suitable acoustic strategy, which can turn a passive playing tool into an active and empathetic music companion.
[0183] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.
[0184] The above is the introduction of the method embodiment, and the scheme described in the present application will be further described through the device embodiment.
[0185] Figure 2 A block diagram of an adaptive sound field device 200 according to an embodiment of the present application is shown, as shown in Figure 2 includes:
[0186] The acquisition module 210 is configured to acquire target data.
[0187] The determining module 220 determines a target sound wave path according to the target data;
[0188] The executing module 230 executes a target playing operation to form a target sound field according to the target sound wave path.
[0189] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the described modules can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0190] Figure 3 A structure diagram of a terminal device or a server suitable for implementing the embodiments of the present application is shown.
[0191] As shown in Figure 3 , the terminal device or the server includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or programs loaded from a storage portion 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the terminal device or the server are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0192] The following components are connected to the I / O interface 305: an input portion 306 including a keyboard, a mouse, and the like; an output portion 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage portion 308 including a hard disk, and the like; and a communication portion 309 including a network interface card such as a LAN card, a modem, and the like. The communication portion 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as necessary. A removable medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 310 as necessary, so that a computer program read therefrom is installed into the storage portion 308 as necessary.
[0193] In particular, according to the embodiments of the present application, the above method flow steps can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product including a computer program carried on a machine-readable medium, which contains program codes for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication portion 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above-described functions defined in the system of the present application are executed.
[0194] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0195] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the aforementioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0196] The units or modules described in the embodiments of the present application can be implemented in the form of software or in the form of hardware. The units or modules described can also be arranged in a processor. In some cases, the names of the units or modules do not constitute a limitation on the units or modules themselves.
[0197] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the methods described in the present application.
[0198] The above description is merely preferred embodiments of the present application and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application described in the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above application concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features applied in the present application (but not limited to) having similar functions.
Claims
1. An adaptive sound field method, characterized in that: include: Get target data; determining a target acoustic wave path according to the target data; According to the target sound wave path, a target playback operation is performed to form a target sound field.
2. The method according to claim 1, characterized in that The target data includes: Target physical environment feature data, target acoustic feature data, and / or target user feature data.
3. The method according to claim 2, characterized in that Determining the target acoustic wave path according to the target data includes: According to the target data, all sound wave paths from each speaker to the user position are calculated to obtain a sound wave path set; Calculating a driving signal vector using an inverse calculation method based on the set of sound wave paths and the number of loudspeakers; Based on the driving signal vector, a target acoustic wave path is determined.
4. The method according to claim 3, characterized in that Calculating all sound wave paths from each speaker to the user position according to the target data includes: According to the target data, all sound wave paths from each speaker to the user position are calculated respectively by ray tracing method; wherein the sound wave paths include reflection paths, diffraction paths, and / or absorption paths.
5. The method according to claim 4, characterized in that The calculating of the driving signal vector by reverse calculation based on the set of sound wave paths and the number of loudspeakers includes: Based on the acoustic wave path set and the number of loudspeakers, the driving signal vector is calculated by reverse calculation using the following formula: X=(H^H H+λI)^{-1}H^H P_target Where H is the transfer matrix, which is set according to the number of sound wave paths and the number of speakers; λ is the regularization coefficient; ^H is the conjugate transpose; P_target is the user target sound pressure in the target user feature data; I is the identity matrix.
6. The method according to claim 5, characterized in that The performing a target playback operation to form a target sound field according to the target sound wave path includes: Inputting the target user feature data into a trained sentiment analysis model to obtain the current user's sentiment category; According to the emotion category and the preset hierarchical intervention strategy, a target playback operation is performed through the target sound wave path to form a target sound field.
7. The method according to claim 6, characterized in that The sentiment analysis model can be trained in the following ways: Generate a training sample set, wherein the training sample includes a script file with annotation information; the annotation information is an emotion label; The sentiment analysis model is trained using samples in the training sample set, with the script file as input and the sentiment label as output. When the unification rate of the output sentiment label and the marked sentiment label meets the preset threshold, the training of the sentiment analysis model is completed.
8. An adaptive sound field device, characterized in that: include: Acquisition module, used to obtain target data; a determination module, which always determines the target acoustic wave path based on the target data; The execution module always performs a target playback operation according to the target sound wave path to form a target sound field.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Hearing protection method and device
CN112577590A
Loudspeaker control method, device and equipment and readable storage medium
CN116055606A
Automobile DSP power amplifier sound effect intelligent processing system
CN119996899A
Personified emotional exchange accompanying type intelligent table lamp
CN120152117A
Sound field controller, sound field control system, and sound field control method
JP2015213249A