Vehicle-mounted sound equipment sound field adjusting method and device, electronic equipment and vehicle
By detecting the position and head relationship of passengers in real time, and combining the sound field adjustment parameters of the vehicle audio system, the speaker output is dynamically adjusted so that the position of the human voice imaging position is directly in front of the passengers. This solves the problem that the sound effect adjustment of the vehicle audio system depends on user operation and achieves the best personalized listening effect.
Patent Information
- Application Number
- CN202511958544.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-02-10
AI Technical Summary
Existing car audio systems rely on subjective user operation for sound effect adjustment, resulting in poor sound quality after adjustment. They cannot dynamically adapt to changes in the position of passengers, especially when the interior space is limited, resulting in poor listening performance.
By detecting the position of the occupants and the relative position of their heads, the imaging position of the human voice in the speaker output sound is dynamically adjusted using pre-stored sound field adjustment parameters, so that it is always located in the area directly in front of the occupants. This is combined with parameters such as output power balance and crossover filtering for precise adjustment.
It achieves optimal sound quality regardless of where passengers are seated, enhancing the immersive listening experience and personalization capabilities, and solving the problems of cumbersome operation and poor sound quality of traditional adjustment modes.
Smart Images

Figure CN121509869A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of vehicles, in particular to a vehicle-mounted sound field adjustment method and device, electronic equipment and vehicle. BACKGROUND
[0002] With the increasing demand of users, the user's listening experience in the vehicle has been upgraded from basic sound effect debugging to scenario-based and personalized immersive experience. For example, the user hopes to simulate the "best C position" listening effect of a concert when listening to music in the vehicle: that is, the vocal elements are always located in front of the listener, and the sound field of musical instruments is distributed on both sides of the stage, thereby forming clear sound image positioning and a stereoscopic listening experience.
[0003] Currently, due to the limitation of the space in the vehicle, the prior art mainly uses sound field adjustment cursor dragging to realize audio adjustment.
[0004] However, this way of the prior art relies on user subjective operation, and the adjusted sound effect is poor. SUMMARY
[0005] In view of the above problems, embodiments of the present application provide a vehicle-mounted sound field adjustment method and device, electronic equipment and vehicle, which are used to solve the problem of poor sound effect adjustment of the vehicle in the prior art.
[0006] According to an aspect of an embodiment of the present application, a vehicle-mounted sound field adjustment method is provided, and the method comprises:
[0007] When it is detected that there is an occupant in the seat area of the vehicle, taking a preset reference point in the seat area as the center, an initial boundary range is determined by using the shoulder width of the occupant;
[0008] Based on the relative position relationship between the head of the occupant and the initial boundary range, a target coordinate value is determined;
[0009] A target sound field adjustment parameter corresponding to the target coordinate value is called from the sound field adjustment parameters pre-stored in the vehicle-mounted power amplifier, and the sound field adjustment parameter comprises at least one of an output power balance parameter and a frequency division filter parameter;
[0010] Based on the target sound field adjustment parameter, the image location of the vocal sound in the loudspeaker output sound is dynamically adjusted until the image location is located in the front area facing the occupant, and the image location refers to the emitting position of the vocal sound heard by the occupant.
[0011] In an optional manner, based on the target sound field adjustment parameter, the image location of the vocal sound in the loudspeaker output sound is dynamically adjusted until the image location is located in the front area facing the occupant, which comprises:
[0012] determining a number of imaged positions of the human voice based on a number of passengers in the vehicle, the number of imaged positions of the human voice being greater than or equal to the number of passengers;
[0013] adjusting dynamically the imaged positions of the human voice in the sound output by each loudspeaker in the vehicle based on the target sound field adjustment parameter and the number of imaged positions, so that there is at least one imaged position of the human voice in a front area facing each passenger.
[0014] In an optional manner, the determining of the target coordinate value based on the relative positional relationship between the head of the passenger and the initial boundary range comprises:
[0015] if the head of the passenger is within the initial boundary range, determining a three-dimensional coordinate value of a representative position point corresponding to the initial boundary range as the target coordinate value;
[0016] if the head of the passenger deviates from the initial boundary range, obtaining a deviation direction and a deviation distance of the head relative to the initial boundary range;
[0017] determining a coordinate value deviation amount based on the deviation direction and the deviation distance;
[0018] determining the target coordinate value based on the coordinate value deviation amount and the three-dimensional coordinate value of the representative position point.
[0019] In an optional manner, the determining of the coordinate value deviation amount based on the deviation direction and the deviation distance comprises:
[0020] establishing a three-dimensional coordinate system with the preset reference point as an origin, a front direction of the vehicle as a longitudinal axis, a left direction of the vehicle or a right direction of the vehicle as a transverse axis, and an upward direction of the vehicle or a downward direction of the vehicle as a vertical axis;
[0021] if the deviation direction is the front direction of the vehicle, determining a longitudinal deviation level according to a distance between the head and the preset reference point on the longitudinal axis;
[0022] if the deviation direction is the left direction of the vehicle or the right direction of the vehicle, determining a transverse deviation level according to a distance between the head and the preset reference point on the transverse axis;
[0023] if the deviation direction is the upward direction of the vehicle or the downward direction of the vehicle, determining a vertical direction deviation level according to a distance between the head and the preset reference point on the vertical axis;
[0024] determining a longitudinal axis coordinate value deviation amount based on the longitudinal deviation level;
[0025] determine a horizontal-axis coordinate value offset amount based on the horizontal offset level;
[0026] determine a vertical-axis coordinate value offset amount based on the vertical offset level.
[0027] In an optional manner, the preset reference point is a center point of the headrest.
[0028] In an optional manner, the method further includes:
[0029] obtaining a user target coordinate value corresponding to a preset position of the user in the vehicle, and at least one candidate sound field adjustment parameter corresponding to the user target coordinate value;
[0030] adjusting a vocal sound in the speaker output sound based on the at least one candidate sound field adjustment parameter, to determine a vocal sound initial image position corresponding to each candidate sound field adjustment parameter;
[0031] determining an initial score corresponding to each candidate sound field adjustment parameter based on a distance between the user target coordinate value and the vocal sound initial image position;
[0032] obtaining a candidate sound field adjustment parameter with the highest initial score as a selected sound field adjustment parameter;
[0033] storing the selected sound field adjustment parameter and the user target coordinate value in the vehicle-mounted power amplifier.
[0034] In an optional manner, the sound field adjustment parameter further includes at least one of a volume adjustment parameter, an equalizer adjustment parameter, a sound wave output delay balance parameter, and a equal loudness compensation parameter.
[0035] According to another aspect of an embodiment of the present application, a vehicle-mounted sound field adjustment device is provided, and the device includes:
[0036] a passenger detection module configured to, when detecting that a seat area of a vehicle has a passenger, determine an initial boundary range with a preset reference point in the seat area as a center and a shoulder width of the passenger as a boundary;
[0037] a position obtaining module configured to determine a target coordinate value based on a relative position relationship between a head of the passenger and the initial boundary range;
[0038] a parameter adjustment module configured to call a target sound field adjustment parameter corresponding to the target coordinate value from sound field adjustment parameters pre-stored in a vehicle-mounted power amplifier, the sound field adjustment parameter including at least one of an output power balance parameter and a frequency division filter parameter;
[0039] An audio adjusting module is configured to dynamically adjust a focus position of a human voice in loudspeaker output sound based on the target sound field adjusting parameter, until the focus position is located in the front area of the occupant.
[0040] According to another aspect of the embodiments of the present application, an electronic device is provided, which comprises a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface are in communication with each other through the communication bus; the memory is configured to store at least one executable instruction, which causes the processor to perform the operations of the method as described above.
[0041] According to still another aspect of the embodiments of the present application, a vehicle is provided, which comprises a vehicle body and the electronic device as described above arranged in the vehicle body.
[0042] The embodiments of the present application capture the relative position relationship between the head of the occupant and the initial boundary range in real time, determine the target coordinate value, and combine the pre-stored sound field adjusting parameter in the vehicle-mounted power amplifier, call the target sound field adjusting parameter corresponding to the target coordinate value, and adjust the focus position of the human voice in the loudspeaker output sound, so that the focus position of the human voice is always located in the front area of the occupant, thereby enabling the occupant to experience the best listening effect.
[0043] The above description is only a summary of the technical solutions of the embodiments of the present application, in order to more clearly understand the technical means of the embodiments of the present application, the embodiments of the present application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the embodiments of the present application more obvious and easy to understand, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS
[0044] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and serve to explain the principles of the application. It should be noted that the various drawings are not to scale, with like reference numerals describing similar features throughout the several views. In the drawings:
[0045] Figure 1 A vehicle-mounted sound field adjusting method flowchart is provided for the embodiments of the present application;
[0046] Figure 2 An adjusting interface diagram is provided for the traditional manual adjusting mode;
[0047] Figure 3 An evaluation score diagram is provided for the embodiments of the present application;
[0048] Figure 4 A human voice focus effect diagram is provided for the embodiments of the present application;
[0049] Figure 5A loudspeaker control flowchart provided for the embodiment of the present application;
[0050] Figure 6 A position information collection flowchart provided for the embodiment of the present application;
[0051] Figure 7 A signal flow diagram provided for the embodiment of the present application;
[0052] Figure 8 A structure diagram of the vehicle-mounted sound system sound field adjustment device provided for the embodiment of the present application;
[0053] Figure 9 A structure diagram of the electronic device provided for the embodiment of the present application. DETAILED DESCRIPTION
[0054] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be embodied in various forms and should not be limited by the embodiments set forth herein.
[0055] With the rapid development of new energy vehicle intelligence, especially in the aspects of cockpit function experience and AI application, the vehicle sound system is a core part of the cockpit and is also the focus of many vehicle manufacturers. From basic sound effect debugging to diversified scene experience, all are expanding around the user's listening experience.
[0056] Researchers found that in real life, when the audience is in the venue to listen to the concert, the home singer is in the center of the stage, and the musical instruments (such as guitar, bass, piano, and drum set) are distributed on both sides of the stage. At this time, the best listening position is the seat opposite the center of the stage at a certain distance from the stage. The audience in this seat (i.e. the best listening position) hears the best listening effect. However, for vehicles, it is a difficult problem to be solved how to make the user achieve the best listening effect similar to listening to a concert in real life when listening to music in the vehicle, due to the space in the vehicle.
[0057] To address the aforementioned issues, this application provides a method, apparatus, electronic device, and vehicle for adjusting the sound field of a vehicle audio system. It pre-stores sound field adjustment parameters corresponding to each position within the vehicle. These parameters ensure that the imaging position of vocals in the sound output from the speaker is directly in front of that position. When an occupant's position changes within the vehicle, the system automatically selects a target sound field parameter corresponding to that real-time position from the pre-stored parameters. Based on this target sound field parameter, the imaging position of vocals in the sound output from the speaker is dynamically adjusted to be directly in front of the occupant, thereby creating an optimal listening experience similar to that of an audience listening to a concert from the best possible position.
[0058] Among them, when a passenger is in a certain position inside the vehicle, if there is a sound image in the area directly in front of their head, and the distance between their head and the image image does not exceed a preset distance threshold, then the passenger's position at this time is the optimal listening position.
[0059] The imaging position of human voices refers to the location from which the voice is perceived by the occupant. In real life, without adjusting the speakers, the location from which the voice is perceived by the occupant is usually the location of the speaker itself. For example, if the speaker is located behind the vehicle's center console screen, then regardless of whether the occupant is sitting in the front passenger seat or in the back seat, the voice they perceive will be coming from behind the screen. The occupant's hearing may be better in the front passenger seat and worse in the back seat; the sound quality varies depending on the seating position.
[0060] Through several experiments, the correspondence between the imaging position of human voices and the sound field adjustment parameters can be established. For example, in order to make the imaging position of human voices in front of the headrest of the driver's seat in the car, the sound field adjustment parameters can be continuously adjusted through experiments until a sound field adjustment parameter X is obtained. This sound field adjustment parameter X can make the imaging position of human voices in the sound output by the car speakers in front of the headrest of the driver's seat in the car. Then, this sound field adjustment parameter X can be associated with the front of the headrest of the driver's seat and stored in the car amplifier.
[0061] If it is necessary to position the vocal imaging position in front of the driver's seat headrest later, this sound field adjustment parameter X can be directly used as the target sound field adjustment parameter.
[0062] This solution ensures that whether the passenger is sitting in the front passenger seat or in the rear seat, the voice they perceive comes from the area directly in front of them, providing the same optimal listening experience regardless of their seating position.
[0063] Figure 1 This is a schematic flowchart of a vehicle audio sound field adjustment method provided in an embodiment of this application. This method can be executed by electronic devices in a vehicle. Figure 1 As shown, the method includes the following steps:
[0064] Step 110: When an occupant is detected in the vehicle's seating area, the initial boundary range is determined using a preset reference point within the seating area as the center and the occupant's shoulder width.
[0065] Step 120: Determine the target coordinates based on the relative position of the occupant's head to the initial boundary range.
[0066] Step 130: Call the target sound field adjustment parameter corresponding to the target coordinate value from the pre-stored sound field adjustment parameters of the vehicle power amplifier.
[0067] The sound field adjustment parameters include at least one of the output power balance parameters and frequency division filtering parameters.
[0068] Step 140: Based on the target sound field adjustment parameters, dynamically adjust the imaging position of the human voice in the loudspeaker output sound until the imaging position is located in the area directly in front of the occupant.
[0069] The imaging position refers to the location from which the voice is perceived by the occupant. Voice imaging means that when an occupant listens to a voice emitted from a speaker, their brain can "see or feel" a clear, concrete, and tangible person emitting the voice from a specific location. By adjusting the imaging position of the voice, the location of the person perceived by the occupant can be changed. For example, if the imaging position of the voice is on the occupant's left, then the occupant will perceive the person singing to their left.
[0070] Traditional car audio system adjustments primarily rely on fixed scene modes and manual adjustment modes to control the sound field distribution. Fixed scene modes include "whole-vehicle equalization" and "driver-priority." "Whole-vehicle equalization" fixes the image of vocals below the windshield, weakening the listening experience for rear passengers. "Driver-priority" only optimizes the sound field around the driver, significantly reducing the listening experience for other passengers. Manual adjustment involves using a sound field adjustment cursor on the central control screen, requiring users to manually drag the cursor to match their position. This method relies on subjective user input, cannot dynamically adapt to real-time changes in passenger position, and requires repeated adjustments for passengers of different heights or sitting postures, resulting in cumbersome operation and a poor listening experience.
[0071] For example, Figure 2 This is a schematic diagram of the adjustment interface in the traditional manual adjustment mode, such as... Figure 2As shown, users can adjust the sound field by dragging the sound field adjustment cursor, and achieve sound field focusing for the driver, passenger, and rear seats.
[0072] In this embodiment, by capturing occupant position information in real time and combining it with the digital signal processing (DSP) of the vehicle audio system, the target sound field parameters corresponding to the occupant position are determined from pre-stored sound field adjustment parameters. This dynamically adjusts the imaging position of the human voice in the speaker output sound, ensuring that the human voice imaging position is always located directly in front of the occupant. This enhances the immersive listening experience and personalization capabilities, solving the problems of cumbersome operation and poor sound quality associated with traditional fixed-scene modes and manual adjustment modes. Furthermore, compared to traditional manual adjustment modes, the vehicle audio sound field method provided in this application does not require manual adjustment by the user. The entire process can be automatically and adaptively adjusted based on the real-time position of the occupant using electronic devices installed in the vehicle.
[0073] Regarding step 110 above, the occupant can refer to the driver in the main driver's seat, or the occupant in the front passenger seat, the rear seat, etc.
[0074] Among these methods, sensors installed inside the vehicle can detect whether there are occupants in the seating area. For example, the Occupant Monitoring System (OMS) can be used to monitor the interior of the vehicle's smart cockpit.
[0075] The preset reference point is a pre-configured location point in the seat area. For example, the center point of the headrest can be selected as the preset reference point. This is mainly because the occupant's head is usually close to the headrest. Selecting the center point of the headrest can make it closer to the occupant's head.
[0076] Using a preset reference point as the center and the occupant's shoulder width as the lateral length of the initial boundary range, combined with the pre-configured longitudinal width and vertical height, a three-dimensional cube region can be constructed as the initial boundary region. This initial boundary region describes the range of movement of the occupant's head when the occupant is sitting in the seat area in a normal posture.
[0077] Regarding step 120 above, after the occupant sits in the seat, their posture may change, especially for rear passengers. Since they do not need to keep their eyes directly forward like the driver, their posture is more varied. For example, a normal posture is with the back of the head against the headrest and eyes looking forward. However, when a rear passenger leans forward to talk to the driver, the back of their head moves away from the headrest, and their body leans forward, thus changing their posture.
[0078] When the occupant's sitting posture changes, only part of the occupant's head may remain within the aforementioned initial boundary range, or it may completely leave the aforementioned initial boundary range and enter outside the initial boundary range.
[0079] Among them, the first type of relative position relationship can be defined as the occupant's head being within the initial boundary range; the second type of relative position relationship is defined as the head being completely outside the initial boundary range; and the third type of relative position relationship is defined as the head being completely within the initial boundary range. The target coordinate values corresponding to these three different relative position relationships can be different.
[0080] For example, when the head is completely within the initial boundary range, the target coordinate value can be the three-dimensional coordinate value of the preset reference point. When the head part is within the initial boundary range, the target coordinate value can be the three-dimensional coordinate value of the preset reference point plus a compensation amount.
[0081] Regarding step 130 above, the car amplifier, also known as the car audio power amplifier, is a core component of the car audio-visual system. It is mainly used to process audio signals to drive speakers.
[0082] For example, in some embodiments, to improve the listening experience, the sound field adjustment parameters may include at least one of the following: crossover filtering parameters, volume adjustment parameters, equalizer adjustment parameters, sound wave output delay balance parameters, output power balance parameters, and equal loudness compensation parameters. The more sound field adjustment parameters invoked, the more precise the audio adjustment.
[0083] In particular, considering that passengers may constantly move their bodies and adjust their sitting posture inside the vehicle, causing the corresponding target coordinate value to change in real time, it is necessary to frequently call the target sound field adjustment parameter corresponding to the target coordinate value from the various pre-stored sound field adjustment parameters.
[0084] In one embodiment, taking the target sound field adjustment parameter as the sound wave output delay balance parameter as an example, the principle is to utilize the time difference between the two ears. Assuming the vehicle has two speakers, a sound is emitted from the left speaker first, and then the sound is emitted from the right speaker a little later. In this case, the sound perceived by the occupant is located slightly to the left.
[0085] At this point, by adjusting the sound wave output delay balance parameter, increasing the delay of the left speaker (or decreasing the delay of the right speaker), the imaging position of the human voice will shift to the right. The more the delay is increased, the further the imaging position of the human voice will shift to the right, thus ensuring that the imaging position of the human voice is in the area directly in front of the occupant.
[0086] In another embodiment, taking the target sound field adjustment parameter as the output power balance parameter as an example, the principle is to utilize the difference in intensity between the two ears. If the loudness of a sound is greater on the left speaker than on the right, then the occupant will perceive the sound as originating from the left.
[0087] At this point, by adjusting the output power balance parameters, increasing the loudness of the left speaker (or decreasing the loudness of the right speaker), the imaging position of the human voice will move to the left. The more the loudness is increased, the further the imaging position of the human voice will move to the left, thus achieving the imaging position of the human voice in the area directly in front of the passenger.
[0088] In another embodiment, taking the target sound field adjustment parameter as the equalizer adjustment parameter as an example, the principle is to utilize the timbre difference and the characteristics of human body's sound perception. The human head, ears and torso will produce different filtering effects on sounds from different directions, thereby changing their frequency response. The brain, through learning, can distinguish the location of the sound source based on the subtle differences in timbre.
[0089] By adjusting the equalizer parameters to appropriately attenuate high frequencies, you can make the vocals sound farther away from the occupants, while boosting the high frequencies will make them sound closer. Additionally, if you want the vocal imaging to be positioned to the left, you can slightly boost the mid-high frequencies (2-5 kHz) of the left speaker by adjusting the equalizer parameters, thus positioning the vocals directly in front of the occupants.
[0090] In another embodiment, taking the target sound field adjustment parameters as the crossover filtering parameters as an example, the principle is to control multiple speakers to ensure that the main mid-frequency range (300 Hz to 3 kHz) of human voices is clearly reproduced by the same speaker, avoiding the problem of the imaging position of human voices diverging due to improper crossover points causing all speakers to participate in the mid-frequency range. For example, when there are multiple speakers in a vehicle, the main mid-frequency range of human voices should be controlled to be emitted by the speaker behind the central control screen as much as possible to adjust the imaging position of human voices and avoid their divergence.
[0091] In another embodiment, taking the target sound field adjustment parameter as the volume adjustment parameter as an example, the principle is to use the volume difference to adjust the imaging position of the human voice. For example, if the volume of the left speaker is large and the volume of the right speaker is small, then the position of the human voice perceived by the passenger may be more to the left. At this time, the volume of the left speaker can be reduced or the volume of the right speaker can be increased, so that the imaging position of the human voice moves to the right, thus realizing that the imaging position of the human voice is in the area directly in front of the passenger.
[0092] In another embodiment, taking the target sound field adjustment parameter as the equal loudness compensation parameter as an example, the principle is that the sensitivity of the human ear to different frequencies changes with the volume. At low volumes, the human ear is not sensitive to low and high frequencies. In low-volume scenarios, equal loudness compensation can automatically boost low and high frequencies, making the position of the vocal image more prominent, so that passengers can more sensitively perceive the specific location of the vocal image.
[0093] Among them, big data training for sound imaging adjustment can be performed to obtain the sound field adjustment parameters corresponding to each position in the car (using the sound field adjustment parameters, the human voice in the sound output by the speaker can be imaged to the area directly in front of that position) and pre-stored, for example, the trained sound field adjustment parameters can be directly implanted into the DSP core of the car amplifier.
[0094] For example, in some embodiments, the sound field adjustment parameters can be trained by the following steps:
[0095] Step (1) Obtain the user target coordinate value corresponding to the user's preset position in the vehicle, and at least one candidate sound field adjustment parameter corresponding to the user target coordinate value;
[0096] Step (2) Based on at least one candidate sound field adjustment parameter, adjust the human voice in the loudspeaker output sound to determine the initial imaging position of the human voice corresponding to each candidate sound field adjustment parameter;
[0097] Step (3) Based on the distance between the user's head coordinates and the initial imaging position of the human voice, determine the initial score corresponding to each candidate sound field adjustment parameter;
[0098] Step (4) Obtain the candidate sound field adjustment parameter with the highest initial score, and use it as the selected sound field adjustment parameter;
[0099] Step (5) save the selected sound field adjustment parameters and user target coordinate values to the vehicle power amplifier.
[0100] In traditional technologies, parameters are usually fixed, which can easily lead to poor adaptability. However, in the embodiments of this application, through a scoring mechanism, the system can automatically learn user preferences (such as preference for low-frequency enhancement) and scene requirements (such as different music genres), dynamically optimize parameters, and achieve more precise sound field adjustment. For example, it can automatically adjust equalizer parameters for different music genres (such as rock and classical), or dynamically optimize gain allocation for different user habits (such as preference for slightly higher volume in the left ear), further optimizing the user's listening experience.
[0101] Regarding step (1) above, the preset location of the user can be either the front seat or the rear seat in the vehicle. The user's target coordinates can be determined by the user's location, posture, and body characteristics (such as height).
[0102] There can be multiple preset positions, such as the driver's seat position, the passenger seat position, the left rear seat position, and the right rear seat position.
[0103] Each preset location has a user target coordinate value, and each user target coordinate value corresponds to at least one candidate sound field adjustment parameter. These candidate sound field adjustment parameters can be obtained through experimental data.
[0104] For example, assuming the user is in the front passenger seat, their target coordinates correspond to three candidate sound field adjustment parameters, namely candidate sound field adjustment parameter A, candidate sound field adjustment parameter B, and candidate sound field adjustment parameter C.
[0105] Adjusting the speaker using candidate sound field adjustment parameter A will position the vocal image to the user's left; using candidate sound field adjustment parameter B will position the vocal image to the user's right; and using candidate sound field adjustment parameter C will position the vocal image directly in front of the user. In other words, different candidate sound field adjustment parameters result in different vocal image positions.
[0106] In addition, the user's target coordinate values can be three-dimensional spatial coordinate values (X, Y, Z). To improve the listening experience, each three-dimensional spatial coordinate value (X, Y, Z) corresponds to at least one candidate sound field adjustment parameter.
[0107] For step (2) above, each candidate sound field adjustment parameter can be used to adjust the sound output by the speaker, and the imaging position of the human voice can be recorded as the initial imaging position of the human voice corresponding to the candidate sound field adjustment parameter.
[0108] Regarding step (3) above, some candidate sound field adjustment parameters may not be what is actually desired and need to be filtered through scoring. For example, using candidate sound field adjustment parameter A to adjust the speaker will make the image position of the human voice in its output sound to be on the left side of the user, which is far away from the user's target coordinate value. However, using candidate sound field adjustment parameter C to adjust the speaker will make the image position of the human voice in its output sound to be directly in front of the user, which is closer to the user's target coordinate value.
[0109] The closer the distance, the higher the initial rating. Ratings can be based on user feedback. For example, Figure 3 A schematic diagram of the evaluation scores provided for embodiments of this application, such as Figure 3 As shown, taking a user sitting in the front passenger seat as an example, the requirement is that the center positioning (i.e., the imaging position of the voice) is located in the area directly in front of the user. The user's specific scoring rules are as follows:
[0110] Aa, the imaging position of the human voice is located in the area directly in front (8-10 points).
[0111] AB, ab: The imaging position of the human voice is close to the area directly in front (5-7 points);
[0112] BC, bc: The imaging position of the human voice is deviated from the area directly in front of the user's seat (3-4 points);
[0113] CD: The imaging position of the vocals is located on the central axis of the vehicle or on the side of the door, which is significantly deviated from the area directly in front of the seat (1-2 points).
[0114] Regarding steps (4) and (5) above, among the three candidate sound field adjustment parameters, the initial score corresponding to candidate sound field adjustment parameter C is the highest. It can be used as the selected sound field adjustment parameter and then associated with the user's target coordinate value and stored in the vehicle power amplifier.
[0115] The next time the vehicle is used, if it detects that a passenger is sitting in the preset position where the user was previously located, and the coordinates of the passenger's head position are the same as the coordinates of the user's head position, then the candidate sound field adjustment parameter C can be directly called as the target sound field adjustment parameter.
[0116] Regarding step 140 above, speakers can be installed in areas such as behind the central control screen and at the bottom of the doors. The speakers in the vehicle can generate multiple imaging positions for human voices inside the vehicle. For example, they can image in front of the headrest of the driver's seat, the headrest of the passenger seat, the headrest of the left rear seat, and the headrest of the right rear seat, generating four imaging positions for human voices.
[0117] When an occupant is seated on one or more seats in the vehicle, the imaging position of the human voice in front of the headrest of that seat can be adjusted accordingly so that it is in the area directly in front of the occupant, thereby enabling the human voice imaging to be located in the area directly in front of each occupant.
[0118] For example, taking the driver as the first to sit in the driver's seat, the imaging position of the voices, which was originally in front of the headrest, can be adjusted to the area directly in front of the driver, thus achieving the driver's optimal listening experience. Furthermore, if a passenger subsequently sits in the front passenger seat, the imaging position of the voices, originally in front of the front passenger's headrest, can be adjusted to the area directly in front of the passenger, thereby achieving the front passenger's optimal listening experience.
[0119] Among them, several sets of sound field adjustment parameters can be pre-configured through debugging or large model training. Each set of sound field adjustment parameters can produce several imaging positions in the car. The imaging positions of human voices produced by different sets of sound field adjustment parameters are different.
[0120] In practical applications, the sound field adjustment parameters can be adjusted through experimental debugging to change the imaging position of human voices.
[0121] For example, using the D group sound field adjustment parameters will cause the audio signal output by the car speakers to generate four imaging positions for human voices, and these four imaging positions are respectively located in front of the headrest of the driver's seat, the headrest of the passenger seat, the headrest of the left rear seat, and the headrest of the right rear seat.
[0122] These pre-configured sound field adjustment parameters and the imaging position of the human voice they generate are stored in the vehicle amplifier, so that after obtaining the real-time location information of the occupants, the appropriate sound field adjustment parameters can be called as the target sound field parameters based on the real-time location information of the occupants.
[0123] Specifically, based on the occupant's real-time location information (specifically, target coordinates), an adaptive algorithm dynamically calls the target sound field adjustment parameters from various sound field adjustment parameters, and adjusts the audio signal output by the speaker accordingly, thereby achieving adaptive adjustment of the imaging position of the human voice, placing it in the area directly in front of the occupant.
[0124] The imaging position of the vocals is located in the area directly in front of each passenger, which allows the passenger to achieve the best listening effect. The best listening effect for the passenger is similar to the listening effect when the main singer is in the center of the stage and the musical instruments (such as guitar, bass, piano, drums) are distributed on both sides of the stage, and the audience sits in a seat directly facing the center of the stage and at a certain distance from the stage.
[0125] In some embodiments, there may be multiple occupants in the vehicle. For example, if there are occupants in all four seats in the vehicle, the optimal listening effect can be achieved for each occupant by adjusting the audio output of the speaker.
[0126] Because a vehicle contains multiple speakers—for example, speakers are located behind the front center console screen and on the left and right rear door panels—the audio signals output from each speaker can be controlled to create multiple imaging positions for human voices. Each imaging position is located in an area directly in front of the passenger, thus satisfying the listening needs of each passenger. The specific steps are as follows:
[0127] Step 1: Determine the number of imaging positions for human voices based on the number of occupants in the vehicle. The number of imaging positions for human voices must be greater than or equal to the number of occupants.
[0128] Step 2: Based on the target sound field adjustment parameters and the number of imaging positions, dynamically adjust the imaging position of human voices in the sound output of each speaker in the vehicle, so that there is at least one human voice imaging position in the area directly in front of each occupant.
[0129] For example, Figure 4 This is a schematic diagram of the human voice imaging effect provided in the embodiments of this application, such as... Figure 4 As shown, the location information of the front and rear occupants of the vehicle can be obtained through the front and rear cabin cameras, respectively. Figure 4 When four occupants are detected inside the vehicle, the system controls the audio signal output of each speaker to ensure that at least four voices are imaged. Furthermore, sound field adjustment parameters are used to adjust the audio signal of each speaker, ensuring that each occupant's directly facing area has at least one voice imaged. This achieves optimal listening quality for each occupant.
[0130] There is no limit to the number of speakers in the vehicle; for example, speakers can be installed in each door and at each seat position.
[0131] In addition, by adjusting the audio signal of the speaker, a corresponding imaging position of the human voice can be pre-generated at each seat. For example, an imaging position of the human voice can be generated in front of the headrest of the driver's seat, in front of the headrest of the passenger seat, in front of the headrest of the left rear seat, and in front of the headrest of the right rear seat.
[0132] Once a passenger sits in a seat in the car, simply adjust the imaging position of the voice corresponding to that seat so that it is in the area directly in front of the passenger.
[0133] For example, if only the driver is in the driver's seat and all other seats are unoccupied, then you only need to adjust the imaging position of the voice pre-generated at the driver's seat so that its imaging position is in the area directly in front of the driver. The imaging positions of the voices corresponding to the other unoccupied seats do not need to be adjusted, which can improve the adjustment efficiency.
[0134] In this embodiment, based on occupant data, the number of imaging positions of the corresponding human voice is determined, and then the imaging position of each human voice is adjusted to the area directly in front of the occupant. This can flexibly adapt to different usage scenarios. When there are many occupants, it can ensure that each occupant can have the best listening effect, while when there are few occupants, it can improve adjustment efficiency.
[0135] The following examples describe in detail how to determine the target coordinate values.
[0136] As mentioned above, when a passenger maintains a normal sitting posture, even if they sway their head from side to side or lean forward slightly, their head will still remain within the initial boundary range. However, if a passenger uses an abnormal sitting posture, such as lying directly across the rear seat, their head will clearly be completely outside this initial boundary range. Based on this, in some embodiments, the target coordinate value can be determined through the following steps:
[0137] Step 11: If the occupant's head is within the initial boundary range, then the three-dimensional coordinates of the representative position point corresponding to the initial boundary range are determined as the target coordinates.
[0138] Step 12: If the occupant's head deviates from the initial boundary range, obtain the offset direction and offset distance of the head relative to the initial boundary range;
[0139] Step 13: Determine the coordinate offset based on the offset direction and deviation distance;
[0140] Step 14: Determine the target coordinates based on the coordinate offset and the three-dimensional coordinates of the representative location point.
[0141] For step 11, the representative location point can refer to the aforementioned preset reference point. When the occupant is within the initial boundary range, the three-dimensional coordinate value of the preset reference point is directly used as the target coordinate value, thereby realizing that the imaging position of the human voice is located in the area directly in front of the occupant.
[0142] For step 12, the offset direction can include left, right, down, up, forward, or backward, and different offset directions have different deviation distances.
[0143] The offset distance can refer to the distance between the occupant's head and the center of the initial boundary range (i.e., the preset reference point).
[0144] For step 13, the coordinate offset can include the horizontal axis offset, the vertical axis offset, and the vertical coordinate offset. For example, if the occupant's head is offset 10 centimeters to the left, the corresponding vertical axis offset is 10 centimeters.
[0145] For step 14, taking the coordinates of the representative position point as (X0, Y0, Z0) and the target coordinates as (X1, Y1, Z1) as an example, multiple offset ranges and corresponding coordinate compensation values for each offset range can be set. When the coordinate value offset falls exactly within the target offset range, the coordinate compensation value corresponding to the target offset range is used to compensate the coordinate value of the representative position point to obtain the target coordinate value.
[0146] For example, when the head shifts to the left, three offset ranges are set: the first offset range is [0, 10], the second offset range is (10, 20], and the third offset range is (20, 30).
[0147] Wherein, the coordinate compensation value corresponding to the first offset range is Y offset1 The coordinate compensation value corresponding to the first offset range is Y. offset2 The coordinate compensation value corresponding to the third offset range is Y. offset3 .
[0148] Continuing with the example of the occupant's head shifting 8 centimeters to the left, the vertical coordinate value Y0 representing the position point plus the coordinate compensation value Y offset1 The target vertical axis coordinate value Y1 is obtained.
[0149] In this embodiment of the application, by obtaining the offset direction and the deviation distance, the coordinate value offset is determined, and the three-dimensional coordinate value of the representative position point is corrected based on the coordinate value offset. This makes the target coordinate value more closely match the actual position of the occupant's head, so as to ensure that the imaging position of the human voice can be in front of the occupant's face, thereby improving the listening effect.
[0150] Furthermore, the coordinate value offset can be determined through the following steps:
[0151] Step 21: Establish a three-dimensional coordinate system with a preset reference point as the origin, the vehicle's forward direction as the orientation of the vertical axis, the vehicle's left or right direction as the orientation of the horizontal axis, and the vehicle's top or bottom as the orientation of the vertical axis.
[0152] Step 22: If the offset direction is the vehicle's forward direction, determine the longitudinal offset level based on the distance between the head and the preset reference point on the longitudinal axis;
[0153] Step 23: If the offset direction is the left or right side of the vehicle, determine the lateral offset level based on the distance between the head and the preset reference point on the horizontal axis;
[0154] Step 24: If the offset direction is above or below the vehicle, determine the vertical offset level based on the distance between the head and the preset reference point on the vertical axis;
[0155] Step 25: Determine the offset of the vertical axis coordinate value based on the vertical offset level;
[0156] Step 26: Determine the offset of the horizontal axis coordinate value based on the horizontal offset level;
[0157] Step 27: Determine the vertical axis coordinate offset based on the vertical offset level.
[0158] In this embodiment, after the occupant sits in the vehicle seat, an X-axis is established with the direction of the occupant's forward head, a Y-axis is established with the direction of the occupant's left side, and a Z-axis is established with the occupant's top. At the same time, a preset reference point is used as the origin to construct a three-dimensional coordinate system. The three-dimensional coordinate system can more accurately describe the occupant's head position and improve the accuracy of the imaging position adjustment of human voice.
[0159] In this embodiment, when the occupant's head shifts, the distance (△X, △Y, △Z) between the occupant's head and the origin of the three-dimensional coordinate system can be collected. Based on △X, the longitudinal offset level is determined; based on △Y, the lateral offset level is determined; and based on △Z, the vertical offset level is determined.
[0160] For example, when the value of △X is greater than 0 and less than or equal to 10, it is a low-level vertical offset; when the value of △X is greater than 10 and less than or equal to 20, it is a medium-level low vertical offset; when the value of △X is greater than 20 and less than or equal to 30, it is a high-level low vertical offset.
[0161] Different offset levels correspond to different coordinate offsets. For example, a low-level vertical offset corresponds to a coordinate offset of 5; a medium-level vertical offset corresponds to a coordinate offset of 10; and a high-level vertical offset corresponds to a coordinate offset of 15.
[0162] In this embodiment of the application, by establishing a three-dimensional coordinate system, the head position is decomposed into components in three directions, making the position description more accurate and consistent. By configuring offset levels and determining the offset amount based on the offset levels, the complex position data processing can be simplified, the computational complexity can be reduced, and the head can be quickly located.
[0163] Furthermore, traditional technologies fail to differentiate and optimize for the listening needs of occupants in different seats, resulting in a significantly weaker listening experience for rear-seat passengers compared to front-seat passengers. Moreover, traditional fixed-parameter modes cannot adapt to dynamic scenarios. In some embodiments, collaborative control of individual speakers can be used to ensure that the vocal image is positioned directly in front of each occupant, guaranteeing optimal listening performance for all occupants. Specifically, Figure 5 This is a schematic diagram of the speaker control process provided in an embodiment of this application, such as... Figure 5 As shown, it includes the following steps:
[0164] Step 510: Obtain the positional relationship between each speaker in the vehicle and the occupants;
[0165] Step 520: Based on the positional relationship of each speaker, determine the audio playback delay parameters and volume gain adjustment parameters for each speaker;
[0166] Step 530: Based on the audio playback delay parameters and volume gain adjustment parameters, coordinate the adjustment of each speaker until the imaging position of the human voice is located in the area directly in front of each passenger.
[0167] In this embodiment, based on audio playback delay parameters and volume gain adjustment parameters, each speaker is controlled in a coordinated manner to adjust the output delay and gain of each speaker. This enables dynamic optimization of the sound field distribution of all speakers in the vehicle, ensuring that the listening experience of passengers in different seats reaches the best effect.
[0168] For example, if there is a left speaker and a right speaker around the passenger, based on the positional relationship, it is found that the passenger is closer to the left speaker and farther from the right speaker. This may cause the image of the human voice to be located to the left of the passenger. In order to adjust the image of the human voice to be in the area directly in front of the passenger, the audio playback delay parameter and volume gain adjustment parameter of the right speaker can be kept unchanged, while the audio playback delay of the left speaker is increased and the volume gain of the left speaker is decreased. This will achieve a balance between the left and right speakers, so that the image of the human voice is in the area directly in front of the passenger.
[0169] Regarding step 510 above, the positional relationship can refer to the distance between the occupant and the speaker. This is because the speakers in the vehicle are located in different positions; for example, some speakers are located on the left side door, while others are located on the right front door. This results in differences in the distance between different speakers and the same occupant. Due to these differences in distance, the occupant may perceive differences in the volume and duration of sound from each speaker.
[0170] For example, speakers include front door speakers, rear door speakers, and A-pillar speakers.
[0171] The speaker's location is fixed. Once the occupant's real-time location information is determined, the distance between them can be determined based on the occupant's real-time location information and the speaker's installation location. For example, the distance between the occupant's head three-dimensional spatial coordinates and the speaker's three-dimensional spatial coordinates can be calculated as the positional relationship between the occupant and the speaker.
[0172] Regarding step 520 above, the audio playback delay parameter can refer to the delay time of the speaker's audio output. For example, for front-seat occupants, since the front door speakers are closer to the front-seat occupants, the audio output of the front door speakers can be delayed.
[0173] Additionally, the volume gain adjustment parameter refers to the speaker volume output level. Continuing with the example of front-seat occupants, since the rear door speakers are farther away from the front-seat occupants, the volume attenuation is greater. In order to balance the volume of the sound heard by the front-seat occupants from the front speakers and the sound from the rear speakers, the volume gain of the rear speakers can be increased to increase the audio output.
[0174] In this embodiment, reasonable audio playback delay parameters and volume gain adjustment parameters can be configured according to the distance between each speaker and the occupant, so that the imaging position of the human voice is located directly in front of the occupant. For example, the closer the speaker is to the occupant, the longer the audio playback delay time; and the farther the speaker is from the occupant, the greater the volume gain.
[0175] In this embodiment, by dynamically adjusting the output parameters of each speaker (including audio playback delay parameters and volume gain adjustment parameters), the system can control the imaging position of human voices in real time in the area directly in front of each passenger, avoiding the problem of weakened hearing due to differences in seat position. For example, the imaging position of human voices for rear passengers can be achieved by increasing the gain of the rear door speakers and adjusting the delay parameters, making their listening experience consistent with that of the front passengers, significantly improving the listening quality of rear passengers, and avoiding the problem of weakened hearing for rear passengers due to being too far from the front speakers in the "whole vehicle equalization" mode.
[0176] To further improve the listening experience, and to ensure that the optimal listening experience dynamically follows changes in the occupant's position, in some embodiments... Figure 6 This is a schematic diagram of the location information collection process provided in the embodiments of this application, such as... Figure 6 As shown, the method may specifically include the following steps:
[0177] Step 610: Check whether there are occupants in each seating area of the vehicle;
[0178] Step 620: If there is an occupant in the seating area, obtain the occupant's head movement range;
[0179] Step 630: Determine the target coordinates based on the head's movable range.
[0180] In this embodiment of the application, facial recognition is used to automatically monitor whether there are occupants in the seat area, which can realize the real-time collection of location information. This allows for the optimal listening effect by dynamically following changes in the occupant's position. Furthermore, by collecting the range of head movement, the occupant's position can be located more accurately. This enables the imaging position of the human voice to be more accurately located in the area directly in front of the occupant's head, thereby further improving the listening effect.
[0181] Regarding step 610 above, the presence of occupants and the driver can be determined by dynamically capturing facial liveness detection using the in-vehicle OMS camera. For example, the front and rear cabin monitoring cameras can be used to perform facial detection on each seat area, obtaining the results. Based on these results, the presence of occupants in each seat area can be determined. For instance, if a face is detected on the front left seat, it is determined that an occupant is present, and that occupant is the driver.
[0182] Regarding step 620, since the occupant's head has auditory organs, by determining the range of head movement, the target sound field adjustment parameters can be called more reasonably, thereby improving the listening effect.
[0183] Among them, the range of head movement varies among different occupants. For example, taller people will have their heads positioned significantly higher than shorter people, and people with wider shoulders may have a wider range of head movement.
[0184] For example, in some embodiments, the headrest can be used as a preset reference point, with the headrest as the center and the occupant's shoulder width as the boundary, thereby determining the range of head movement.
[0185] In this embodiment, by using shoulder width as a standard to measure the range of head movement, the vehicle can more accurately locate the real-time position of the occupants, distinguish people with different characteristics (such as height), accurately locate the imaging position requirements of each different occupant for the human voice, eliminate problems such as uneven hearing between the left and right ears and sound field deviation, simulate the "best listening position" effect of a concert, so that occupants, whether in the front or back row, can feel that the imaging position of the human voice is in the area directly in front of them, significantly improving the immersive listening experience.
[0186] Regarding step 630 above, in some embodiments, at least one location point can be extracted from the movable range of the head, and the three-dimensional spatial coordinate value of the location point can be collected and determined as the target coordinate value.
[0187] The more location points there are, the more precise the control becomes, which can further improve the listening experience.
[0188] Specifically, the three-dimensional spatial coordinates (X, Y, Z) of each position point within the movable range of the head can be extracted as the real-time position of the occupant.
[0189] Each three-dimensional spatial coordinate value (X, Y, Z) will automatically parse a corresponding parameter call signal to call the corresponding target sound field adjustment parameters.
[0190] Additionally, if the head has a large range of motion, the three-dimensional spatial coordinates of N representative locations can be selectively extracted to serve as the occupant's real-time position. For example, several locations on both sides of the head near the ears can be selected to extract their three-dimensional spatial coordinates.
[0191] This application also provides a vehicle audio sound field system, which may include an OMS in-cabin camera, a cockpit domain controller (CDC), a front-end display screen, and a vehicle power amplifier assembly.
[0192] The system incorporates a DSP chip within the vehicle's power amplifier assembly to train and store large-scale data for sound imaging adjustment. This data is then combined with real-time facial liveness detection using front and rear cabin monitoring cameras. Based on the coordinates (X, Y, Z) of the human head's movement (e.g., for a rear passenger sitting in the right rear seat, the OMS collects the corresponding three-dimensional coordinates of the head within the headrest's range of movement, using a 75% human body model), the data is sent back to the cabin domain controller for analysis. Each coordinate value automatically corresponds to a Controller Area Network (CAN) signal, which is then sent to the vehicle's power amplifier assembly. The power amplifier assembly then calls built-in audio effect-related parameters (e.g., 0X1 represents the head coordinates of passenger 1, and so on) to automatically adjust the vocal imaging in the music to the area directly in front of the human body, ensuring that the passenger is always in the optimal listening position and that the sound experience is not affected by changes in the position of the ears.
[0193] The more head coordinates sampled, the more precise the control. By using an OMS camera to dynamically capture the faces of passengers and drivers for liveness detection, head coordinates are obtained. This ensures that the sound experience is not affected by changes in head position due to different body heights, solving the problem that traditional technologies cannot adapt to changes in passenger height to improve the listening experience.
[0194] The sound field adjustment parameters can be adjusted to cover the speakers in each door and each passenger seat. The more sound adjustment parameters there are, the more head sampling coordinates are sampled, the more precise the sound effect adjustment will be, and the better the final listening effect will be.
[0195] Figure 7 This is a schematic diagram of signal flow provided in the embodiments of this application, such as... Figure 7 As shown, the image signal of the head position monitored by the OMS camera is transmitted back to the cockpit domain controller in real time. The cockpit domain controller converts it into a control signal and sends it to the vehicle power amplifier assembly. When the vehicle power amplifier assembly receives the control signal, it automatically calls the sound field adaptive algorithm built into the DSP chip to perform the following steps:
[0196] Step 1: Play music. The multimedia is transmitted from the cockpit domain controller to the vehicle amplifier assembly via A2B. At this time, the cockpit's default sound field is adjusted to the whole vehicle equalization mode, that is, the main sound field heard by each seat is concentrated in the middle position below the windshield of the front row of the cockpit.
[0197] Step 2: The user turns on this switch via the central control switch (on & off adaptive sound field) on the front display screen. At this time, OMS intervenes to monitor and transmits the coordinates of the head images of each seat to the cockpit domain controller in real time. After receiving the image coordinate signal, the cockpit domain controller converts it into a CAN control signal and sends it to the vehicle power amplifier assembly for adaptive sound field output control, adjusting the sound field to the area directly in front of each user.
[0198] Step 3: The position of the vocal imaging is controlled in the area directly in front of each seated person. The specific implementation process is as follows:
[0199] Step 3.1: OMS collects the coordinates of the human head position and transmits the coordinate signal back to the cockpit domain controller.
[0200] Step 3.2: The cockpit domain controller converts the coordinate signal into a CAN signal and sends it to the vehicle power amplifier assembly.
[0201] Step 3.3: The vehicle power amplifier assembly receives the corresponding CAN signal and automatically retrieves the sound effect parameters from the DSP chip.
[0202] The DSP chip's built-in sound effect parameters are not fixed. Instead, they are dynamically adjusted by the vehicle amplifier assembly using an adaptive algorithm based on the coordinate positions collected by the OMS. Coordinates of the head position for each seat and for each person's height need to be collected, and all speakers in the vehicle need to participate in sound production. Extension and gain adjustments are made through the door panel speakers corresponding to each seat, and each door speaker participates in sound image compensation. For example, the main sound field position is controlled in real time in the area directly in front of each person in each seat to achieve the best stage feel and the best listening experience. This prevents issues such as different sound imaging, sound volume, or deviations to the left, right, top, or bottom caused by different sound field focusing positions within the cabin, thus avoiding a degraded listening experience for the user.
[0203] In this embodiment, the face position signal detected by the OMS camera received by the cockpit domain controller is fed back to the cockpit domain controller. The dynamic signal of the cockpit domain controller is sent to the vehicle power amplifier via CAN. The vehicle power amplifier identifies and calls the corresponding sound field tuning parameters of the DSP in the vehicle power amplifier according to the control signal of the head position sent by the cockpit domain controller, so as to achieve dynamic adaptive matching.
[0204] In addition, since the vehicle audio sound field system requires the head position signal collected by the camera to be transmitted back to the cockpit domain controller, and the cockpit domain controller to send the signal transmitted back by the camera to the vehicle power amplifier via the CAN network, the final actuator that presents accurate sound field positioning is the vehicle power amplifier assembly. Therefore, the communication matrix between the cockpit domain controller and the vehicle power amplifier needs to be adaptively programmed and controlled using CAN signals. The specific CAN signals are monitored in real time by the head position image recognized by the camera and uploaded to the cockpit domain controller in real time, and the cockpit domain controller sends the signal to the vehicle power amplifier.
[0205] Figure 8 This is a schematic diagram of the structure of the vehicle audio sound field adjustment device provided in an embodiment of this application. Figure 8 As shown, the vehicle audio sound field adjustment device 800 includes:
[0206] The occupant detection module 810 is used to determine the initial boundary range by using a preset reference point in the seat area as the center and the shoulder width of the occupant when an occupant is detected in the seat area of the vehicle.
[0207] The position acquisition module 820 is used to determine the target coordinate value based on the relative positional relationship between the occupant's head and the initial boundary range;
[0208] The parameter adjustment module 830 is used to call the target sound field adjustment parameter corresponding to the target coordinate value from the sound field adjustment parameters pre-stored in the vehicle power amplifier. The sound field adjustment parameter includes at least one of the output power balance parameter and the frequency crossover filter parameter.
[0209] The audio adjustment module 840 is used to dynamically adjust the imaging position of human voice in the speaker output sound based on the target sound field adjustment parameters, until the imaging position is located in the area directly in front of the occupant. The imaging position refers to the position where the human voice is emitted when heard by the occupant.
[0210] In one alternative approach, the audio adjustment module can specifically be used for:
[0211] The number of imaging positions of human voices is determined based on the number of occupants in the vehicle, and the number of imaging positions of human voices is greater than or equal to the number of occupants.
[0212] Based on the target sound field adjustment parameters and the number of imaging positions, the imaging position of human voice in the sound output of each speaker in the vehicle is dynamically adjusted so that there is at least one human voice imaging position in the area directly in front of each occupant.
[0213] In one alternative approach, the location acquisition module can specifically be used for:
[0214] If the occupant's head is within the initial boundary range, then the three-dimensional coordinates of the representative position point corresponding to the initial boundary range will be determined as the target coordinates.
[0215] If the occupant's head deviates from the initial boundary range, then obtain the offset direction and offset distance of the head relative to the initial boundary range;
[0216] The coordinate value offset is determined based on the offset direction and deviation distance;
[0217] The target coordinates are determined based on the coordinate offset and the three-dimensional coordinates representing the location point.
[0218] In one alternative approach, the location acquisition module can specifically be used for:
[0219] A three-dimensional coordinate system is established with a preset reference point as the origin, the vehicle's forward direction as the orientation of the vertical axis, the vehicle's left or right direction as the orientation of the horizontal axis, and the vehicle's top or bottom as the orientation of the vertical axis.
[0220] If the offset direction is the vehicle's forward direction, the longitudinal offset level is determined based on the distance between the head and the preset reference point on the longitudinal axis;
[0221] If the offset direction is to the left or right of the vehicle, the lateral offset level is determined based on the distance between the head and the preset reference point on the horizontal axis.
[0222] If the offset direction is above or below the vehicle, the vertical offset level is determined based on the distance between the head and the preset reference point on the vertical axis.
[0223] Determine the offset of the vertical axis coordinate value based on the vertical offset level;
[0224] Determine the offset of the horizontal axis coordinate value based on the horizontal offset level;
[0225] The offset of the vertical axis coordinate value is determined based on the vertical offset level.
[0226] In one optional approach, the preset reference point is the center point of the headrest.
[0227] In an alternative approach, a parameter pre-storage module is also included, used for:
[0228] Obtain the user target coordinate value corresponding to the user's preset position in the vehicle, and at least one candidate sound field adjustment parameter corresponding to the user target coordinate value;
[0229] Based on at least one candidate sound field adjustment parameter, the human voice in the loudspeaker output sound is adjusted to determine the initial imaging position of the human voice corresponding to each candidate sound field adjustment parameter;
[0230] Based on the distance between the user's target coordinates and the initial imaging position of the human voice, the initial score corresponding to each candidate sound field adjustment parameter is determined;
[0231] Obtain the candidate sound field adjustment parameter with the highest initial score, and use it as the selected sound field adjustment parameter;
[0232] The selected sound field adjustment parameters and user target coordinate values are pre-stored to the vehicle amplifier.
[0233] In one alternative approach, the sound field adjustment parameters may also include at least one of volume adjustment parameters, equalizer adjustment parameters, sound wave output delay balance parameters, and equal loudness compensation parameters.
[0234] Compared to traditional in-vehicle audio adjustment, the in-vehicle audio sound field device provided in this application can adjust the position of the human voice imaging without the need for manual adjustment by the user, and can adaptively adjust the position of the human voice imaging based on the real-time position of the occupants.
[0235] The vehicle audio sound field device provided in this application captures the position information of passengers in real time and dynamically adjusts the digital signal processing parameters of the vehicle audio system to ensure that the imaging position of human voices is always located in the area directly in front of each passenger, thereby enhancing the immersive listening experience and personalized adaptation capabilities, and solving the problems of cumbersome operation and poor listening effect of traditional fixed scene mode and manual adjustment mode.
[0236] Figure 9 The diagram provided is a structural illustration of an electronic device according to an embodiment of this application. The specific embodiments of this invention do not limit the specific implementation of the electronic device. Figure 9As shown, the electronic device may include: one or more processors 901 and a communication interface 903; the processor 901 is used to perform the steps in the above method embodiments.
[0237] The electronic device may also include a memory 902 and a communication bus 904.
[0238] The processor 901, communication interface 903, and memory 902 communicate with each other via communication bus 904. Communication interface 903 is used for communication with other network elements such as clients or other servers. The processor 901 executes program 905, specifically performing the relevant steps in the above method embodiments.
[0239] Specifically, program 905 may include program code comprising computer-executable instructions. Processor 901 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The vehicle may include one or more processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0240] Memory 902 is used to store program 905. Memory 902 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0241] Specifically, program 905 can be called by processor 901 to cause the electronic device to perform the following operations:
[0242] The real-time location information of each occupant in the vehicle is obtained, including the three-dimensional spatial coordinates of at least one point within the range of motion of the occupant's head.
[0243] The target sound field adjustment parameter corresponding to the three-dimensional spatial coordinates of at least one position point is called from the pre-stored sound field adjustment parameters of the vehicle power amplifier. The sound field adjustment parameters include at least one of the output power balance parameters and the frequency crossover filter parameters.
[0244] Based on the target sound field adjustment parameters, the imaging position of human voice in the loudspeaker output sound is dynamically adjusted until the imaging position is located in the area directly in front of each passenger. The imaging position refers to the position where the human voice is emitted and heard by the passenger.
[0245] The electronic device provided in this application captures occupant location information in real time and dynamically adjusts the digital signal processing parameters of the vehicle audio system to ensure that the human voice image is always located in the area directly in front of each occupant, thereby enhancing the immersive listening experience and personalized adaptation capabilities, and solving the problems of cumbersome operation and poor sound effect of traditional fixed scene mode and manual adjustment mode.
[0246] Furthermore, this application also provides a vehicle, which includes a vehicle body and electronic devices as described above disposed in the vehicle body.
[0247] This invention provides a computer-readable storage medium storing at least one executable instruction that, when executed on a vehicle or data monitoring device, causes the vehicle or data monitoring device to perform the method described in any of the above-described method embodiments. The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Furthermore, this invention is not directed to any particular programming language.
[0248] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Furthermore, the embodiments of this invention are not directed to any particular programming language.
[0249] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. Similarly, for the sake of brevity and to aid in understanding one or more aspects of the invention, in the description of exemplary embodiments of the invention above, various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof. The claims, which follow the detailed description, are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.
[0250] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components, except that at least some of such features and / or processes or units are mutually exclusive.
[0251] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. A method for adjusting the sound field of a car audio system, characterized in that, The method includes: When an occupant is detected in the vehicle's seating area, an initial boundary range is determined using a preset reference point within the seating area as the center and the occupant's shoulder width. The target coordinates are determined based on the relative position of the occupant's head and the initial boundary range. The target sound field adjustment parameter corresponding to the target coordinate value is called from the pre-stored sound field adjustment parameters of the vehicle power amplifier. The sound field adjustment parameter includes at least one of the output power balance parameter and the frequency crossover filter parameter. Based on the target sound field adjustment parameters, the imaging position of the human voice in the loudspeaker output sound is dynamically adjusted until the imaging position is located in the area directly in front of the occupant. The imaging position refers to the position where the human voice is heard by the occupant.
2. The method according to claim 1, characterized in that, The step of dynamically adjusting the imaging position of human voices in the loudspeaker output sound based on the target sound field adjustment parameters until the imaging position is located in the area directly in front of the occupant includes: Based on the number of occupants in the vehicle, the number of imaging positions of the human voice is determined, wherein the number of imaging positions of the human voice is greater than or equal to the number of occupants. Based on the target sound field adjustment parameters and the number of imaging positions, the imaging position of human voice in the sound output by each speaker in the vehicle is dynamically adjusted so that there is at least one human voice imaging position in the area directly in front of each occupant.
3. The method according to claim 1, characterized in that, Determining the target coordinates based on the relative positional relationship between the occupant's head and the initial boundary range includes: If the occupant's head is within the initial boundary range, then the three-dimensional coordinates of the representative position point corresponding to the initial boundary range are determined as the target coordinates. If the occupant's head deviates from the initial boundary range, then the offset direction and offset distance of the head relative to the initial boundary range are obtained; Based on the offset direction and the deviation distance, determine the coordinate value offset; The target coordinate value is determined based on the coordinate value offset and the three-dimensional coordinate value of the representative location point.
4. The method according to claim 3, characterized in that, Determining the coordinate value offset based on the offset direction and the deviation distance includes: A three-dimensional coordinate system is established with the preset reference point as the origin, the vehicle's forward direction as the orientation of the vertical axis, the vehicle's left or right side as the orientation of the horizontal axis, and the vehicle's top or bottom as the orientation of the vertical axis. If the offset direction is the vehicle's forward direction, then the longitudinal offset level is determined based on the distance between the head and the preset reference point on the longitudinal axis; If the offset direction is the left or right side of the vehicle, the lateral offset level is determined based on the distance between the head and the preset reference point on the horizontal axis. If the offset direction is above or below the vehicle, the vertical offset level is determined based on the distance between the head and the preset reference point on the vertical axis. Based on the aforementioned vertical offset level, determine the offset of the vertical axis coordinate value; Based on the aforementioned horizontal offset level, determine the horizontal axis coordinate value offset; Based on the vertical offset level, the vertical axis coordinate value offset is determined.
5. The method according to any one of claims 1-4, characterized in that, The preset reference point is the center point of the headrest in the seating area.
6. The method according to claim 1, characterized in that, The method further includes: Obtain the user target coordinate value corresponding to the user's preset position in the vehicle, and at least one candidate sound field adjustment parameter corresponding to the user target coordinate value; Based on the at least one candidate sound field adjustment parameter, the human voice in the loudspeaker output sound is adjusted to determine the initial imaging position of the human voice corresponding to each candidate sound field adjustment parameter; Based on the distance between the user target coordinates and the initial imaging position of the human voice, an initial score is determined for each candidate sound field adjustment parameter. Obtain the candidate sound field adjustment parameter with the highest initial score, and use it as the selected sound field adjustment parameter; The selected sound field adjustment parameters and the user target coordinate values are pre-stored in the vehicle power amplifier.
7. The method according to any one of claims 1-5, characterized in that, The sound field adjustment parameters also include at least one of the following: volume adjustment parameters, equalizer adjustment parameters, sound wave output delay balance parameters, and equal loudness compensation parameters.
8. A vehicle audio sound field adjustment device, characterized in that, The device includes: The occupant detection module is used to determine the initial boundary range by using a preset reference point in the seat area as the center and the shoulder width of the occupant when an occupant is detected in the seat area of the vehicle. The location acquisition module is used to determine the target coordinate value based on the relative positional relationship between the occupant's head and the initial boundary range; The parameter adjustment module is used to call the target sound field adjustment parameter corresponding to the target coordinate value from the sound field adjustment parameters pre-stored in the vehicle power amplifier. The sound field adjustment parameter includes at least one of the output power balance parameter and the frequency crossover filter parameter. The audio adjustment module is used to dynamically adjust the imaging position of the human voice in the speaker output sound based on the target sound field adjustment parameters, until the imaging position is located in the area directly in front of the occupant. The imaging position refers to the position where the human voice is heard by the occupant.
9. An electronic device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation of the method as described in any one of claims 1-7.
10. A vehicle, characterized in that, It includes a vehicle body and an electronic device as described in claim 9 disposed in the vehicle body.